What The Role Is
HTX is Singapore's Science and Technology agency that brings together diverse scientific and engineering capabilities to develop transformative, operationally ready solutions for public safety. As a statutory board under the Ministry of Home Affairs, HTX works at the forefront of science and technology to empower the Home Team with cutting-edge capabilities. Guided by our mission to amplify, augment and accelerate the Home Team's advantage, we are committed to keeping Singapore the safest place on planet Earth.
xCloud is dedicated to elevating the enterprise experience through cutting-edge cloud capabilities, including:
- Sustainable enterprise data centres
- Hybrid cloud platforms
- Advanced cloud security measures
- Integrated Development, Security, and Operations (DevSecOps)
- Innovative enterprise Software-as-a-Service (SaaS) solutions
- AI-powered enterprise applications
As Head, MLOps, you will lead a team of engineers in designing, developing, and maintaining secure end-to-end DevSecOps solutions and processes. This is a hands-on technical leadership role sitting at the intersection of systems engineering, applied ML, and enterprise governance — pivotal in ensuring the security, scalability, and operational efficiency of software development and engineering workflows across HTX. You will lead the design, architecture, and delivery of HTX's enterprise-scale AI platform, spanning the inference stack that serves in-house and open-source LLMs, and the agentic AI platform that enables officers across Home Team Departments to build, govern, and automate their own workflows.
What You Will Be Working On
- AI Platform Architecture & Delivery: Define and evolve the reference architecture for an enterprise agentic AI platform covering agent harnesses, sandboxing strategy, MCP connector gateways, durable execution, skills libraries, approval workflows, and structured audit trails.
- Shift the organisation from bespoke per-workflow builds to a platform where officers can describe what they need in natural language and have governed, versioned, repeatable workflows generated and executed.
- Inference Infrastructure: Own the roadmap and production operation of the LLM inference stack, covering vLLM/SGLang/TGI-class serving, GPU utilisation, quantisation, speculative decoding, KV-cache optimisation, and related throughput and latency work. Ensure the stack runs reliably across on-prem GPU clusters and air-gapped environments.
- Senior Leadership Engagement: Shape and communicate the AI platform strategy to chief executives and senior leadership across the Home Team. This includes building the business case, securing funding, and translating rapidly evolving technical trends (agentic harnesses, MCP adoption, managed agent services) into decisions the organisation can act on.
- Team Building & Mentorship: Grow the engineering team through hiring, onboarding, and structured internship programmes.
- Define reading lists, career paths, and technical standards. Mentor engineers across inference, platform, and agentic workflow areas
What We Are Looking For
- AI Platform Architecture & Delivery: Define and evolve the reference architecture for an enterprise agentic AI platform covering agent harnesses, sandboxing strategy, MCP connector gateways, durable execution, skills libraries, approval workflows, and structured audit trails.
- Shift the organisation from bespoke per-workflow builds to a platform where officers can describe what they need in natural language and have governed, versioned, repeatable workflows generated and executed.
- Inference Infrastructure: Own the roadmap and production operation of the LLM inference stack, covering vLLM/SGLang/TGI-class serving, GPU utilisation, quantisation, speculative decoding, KV-cache optimisation, and related throughput and latency work. Ensure the stack runs reliably across on-prem GPU clusters and air-gapped environments.
- Senior Leadership Engagement: Shape and communicate the AI platform strategy to chief executives and senior leadership across the Home Team. This includes building the business case, securing funding, and translating rapidly evolving technical trends (agentic harnesses, MCP adoption, managed agent services) into decisions the organisation can act on.
- Team Building & Mentorship: Grow the engineering team through hiring, onboarding, and structured internship programmes.
- Define reading lists, career paths, and technical standards. Mentor engineers across inference, platform, and agentic workflow areas
All new hires are appointed on a two-year contract in the first instance and will be assessed and considered for permanent tenure over time, based on performance.
As part of the shortlisting process for this role, you may be required to complete a medical declaration and/or undergo further assessment.
All applicants will be updated on the status of their applications within 4 weeks upon closing of the advertisement.