We are seeking a hands-on Platform Lead to head our MLOps infrastructure and lead a high-performing DevOps team. Collaborating closely with Data Science and Engineering leaders, you will define our AI/ML infrastructure roadmap, automate multi-cloud provisioning via Infrastructure as Code (IaC), and deploy Large Language Models (LLMs) into production. By architecting resilient LLMOps pipelines and utilising serving frameworks like Triton, vLLM, or Hugging Face TGI, you will guarantee ultra-low latency, high availability, and proactive model monitoring.
Crucially, you will instil financial accountability by establishing robust FinOps frameworks to manage high-cost GPU/CPU cloud budgets. Through auto-scaling, strategic spot instances, and down-scaling policies, you will eliminate compute waste and provide complete visibility into the unit economics of training and serving LLMs. Alongside cost-optimisation, you will enforce rigorous data governance, robust platform security, and 24/7 incident response, ensuring our AI initiatives scale securely and sustainably.