NCS Group

#EG Senior / LLMOps Engineer

NCS Group  •  Singapore, SG (Onsite)  •  2 days ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

NCS is a leading AI Tech Services company. With a 15,000-strong team across the Asia Pacific, NCS scales its platforms and capabilities to provide clients with greater agility and AI expertise across a range of Industries. Embracing a strong ecosystem of global partners, NCS transforms technology services delivery combining AI with digital resilience to drive real business impact. NCS is a subsidiary of the Singtel Group.

This role sits within NCS AI Central's (AIC) Forward Deployed Engineering (FDE) model — the combined capability that takes AI solutions from proof-of-concept through to hardened production systems. You will operate across both fast-moving FDE engagements (POC/POV, pilot deployments for strategic and lighthouse clients) and steady-state system development and maintenance work — bringing the same rigor and a reusable, asset-fed approach to both.

What will you do?

1. Model & Deployment Operations

  • Own the deployment pipeline for LLM and agentic applications — versioning models, prompts, and configurations as code, with safe rollout and rollback paths.
  • Build repeatable CI/CD pipelines for AI services (containerised, on Kubernetes) so new model or prompt versions ship without manual intervention.
  • Support rapid, disposable environment spin-up for FDE POC/POV work, then harden the same pipeline into a production-grade deployment when an engagement scales.

2. Monitoring & Reliability

  • Instrument LLM applications with observability for latency, error rates, output drift, and hallucination signals — not just infrastructure uptime.
  • Define and track SLAs/SLOs for production AI services in partnership with AI Architects and Solution Architects.
  • Set up alerting and runbooks so issues in live AI systems are caught and triaged before clients notice.

3. Cost & Performance Management

  • Monitor token spend and inference cost per model/engagement; flag anomalies and right-sizing opportunities in partnership with AI FinOps.
  • Tune routing between model tiers (frontier vs. smaller/fine-tuned models) for cost-performance balance at production scale.

4. FDE & Development/Maintenance Coverage

  • During FDE engagements: stand up lightweight, reusable deployment scaffolding that lets AI Engineers iterate quickly on POC/POV without operational overhead.
  • During system development & maintenance engagements: take ownership of steady-state production operations, patching, upgrades, and incident response for live AI systems.
  • Contribute reusable deployment patterns back into the shared internal asset library so future engagements start from a hardened baseline, not from zero.

5. Collaboration

  • Work closely with AI Engineers, Cloud Architects, and AI Site-facing leads to ensure a smooth handoff from POC to Scale to Operate.
  • Mentor engineers on LLMOps practices and participate in PRR (Production Readiness Review) gate reviews.

Role Levels We Are Hiring For

We are hiring at two levels for this role. All responsibilities above apply to both; the distinction is in scope of ownership, years of experience, and seniority of judgement expected.

LLMOps Engineer

  • 2–4 years of relevant experience. Executes deployment pipelines, monitoring setup, and cost tracking for individual engagements, under guidance from a Senior LLMOps Engineer or AI Architect.
  • Builds and maintains CI/CD and observability for one or two engagements at a time; escalates novel production incidents to senior team members.

Senior LLMOps Engineer

  • 5+ years of relevant experience, including prior ownership of production AI/ML systems end-to-end. Sets LLMOps standards and reusable deployment patterns across multiple engagements.
  • Leads incident response for critical production issues, mentors junior LLMOps and AI Engineers, and engages directly with client technical stakeholders on production-readiness and reliability.

Qualifications

The ideal candidate should possess:

  • 2+ years in DevOps/MLOps/platform engineering, with hands-on exposure to LLM or ML systems in production (not just prototypes); 5+ years with end-to-end ownership expected at Senior level.
  • Hands-on with containers, Kubernetes, CI/CD (GitHub Actions/GitLab/Jenkins), and Infrastructure-as-Code (Terraform).
  • Practical experience deploying and operating LLM applications (RAG/agentic systems) at scale, including model gateway/routing patterns.
  • Strong scripting/programming ability (Python and/or Go), comfortable working across cloud platforms (AWS/Azure/GCP; GCC/HCC exposure a plus).
  • Working knowledge of observability tooling (OpenTelemetry, Prometheus/Grafana, ELK/OpenSearch) applied to AI-specific signals (drift, hallucination rate, token cost).
  • Comfortable operating in both fast-paced, ambiguous POC/POV settings and disciplined, SLA-driven production support environments.
  • Working knowledge of the China AI model/tech stack (e.g., DeepSeek, Qwen, GLM, Kimi, MiniMax) — deployment patterns, licensing, and self-hosting requirements.

Preferred Qualifications

  • Experience with model gateways/routers (LiteLLM, Bedrock, Vertex, Azure OpenAI) and vector databases (pgvector, Pinecone, Weaviate).
  • Exposure to regulated or government environments (IM8/VAPT, PDPA) and multi-tenant/data-residency patterns.
  • Familiarity with agent orchestration frameworks (LangGraph, Semantic Kernel) and prompt/version control tooling.
  • Experience contributing to or maintaining a reusable internal platform/asset library.
  • Hands-on deployment/self-hosting experience with Chinese open-weight models (DeepSeek, Qwen, GLM) via vLLM/TGI or similar inference runtimes.

Tech Stack (Illustrative)

  • Languages: Python, Go/TypeScript
  • Platform: Docker, Kubernetes, Helm, Argo CD, Terraform, Vault
  • LLM Runtime: OpenAI/Azure OpenAI/Bedrock/Vertex; DeepSeek/Qwen/GLM (China stack); vLLM/TGI; model gateways/routers
  • Observability: OpenTelemetry, Prometheus/Grafana, ELK/OpenSearch, cost meters per request/model
  • Storage/Search: Postgres, Redis; pgvector/Pinecone/Weaviate

Additional Information

Why Join NCS?

Grow with Us

  • Work on cutting-edge AI products that shape the future of technology
  • Collaborate with talented, passionate teams across research, engineering, and design
  • Access continuous learning opportunities and career development pathways

Make an Impact

  • Transform AI research into products that solve real problems for clients and users
  • Drive innovation in a leading Technology Services Firm with regional presence
  • Contribute to building a better future through responsible, human-centred AI

Thrive in Our Culture

  • Experience a human-to-human approach where relationships and collaboration matter
  • Be part of Team NCS, where bold ideas meet practical execution
  • Enjoy a supportive environment that values diversity, inclusion, and respect

We are driven by our AEIOU beliefs—Adventure, Excellence, Integrity, Ownership, and Unity—and we seek individuals who embody these values in both their professional and personal lives. We are committed to our Impact: Valuing our clients, Growing our people, and Creating our future.

Together, we make the extraordinary happen.

Learn more about us at ncs.co and visit our LinkedIn career site.

Scam Alert

We are aware of fraudulent job offers and impersonations of NCS recruiters. Phishing emails using convincing-looking but fake addresses are also commonly used to trick you into thinking that they come from official NCS sources.

Please note that all official communications from NCS Group will only be sent from verified corporate email addresses. Always check that the sender’s email address ends with the genuine NCS domain, @ncs.com.sg and beware of extra letters, symbols or misspellings. When in doubt, verify the sender’s identity by contacting us at reachus@ncs.com.sg.

NCS Group

About NCS Group

NCS, a subsidiary of Singtel Group, is a leading technology services firm with presence in Asia Pacific and partners with governments and enterprises to advance communities through technology. Combining the experience and expertise of its 13,000-strong team across 56 specialisations, NCS provides differentiated and end-to-end technology services to clients with its NEXT capabilities in digital, data, cloud and platforms, as well as core offerings in application, infrastructure, engineering and cybersecurity. NCS also believes in building a strong partner ecosystem with leading technology players, research institutions and start-ups to support open innovation and co-creation. For more information, visit ncs.co.

Industry
IT & Software
Company Size
10,000+ employees
Headquarters
Singapore, SG
Year Founded
1981
Website
ncs.co
Social Media