Role Overview
We are looking for an experienced MLOps Engineer to help build, deploy and scale machine learning solutions across the organisation. This role sits at the intersection of Machine Learning, Data Engineering, Cloud and DevOps, ensuring models can move reliably from development into production and operate effectively at scale.
The ideal candidate will have a strong engineering mindset, hands-on cloud and automation experience, and a passion for building robust platforms that enable Data Scientists and ML Engineers to deliver AI solutions faster and more reliably.
Key Responsibilities
- Build and maintain end-to-end MLOps pipelines for developing, testing, deploying and monitoring machine learning models.
- Automate ML workflows across the model lifecycle, from experimentation through to production.
- Develop CI/CD pipelines for machine learning models, data pipelines and supporting applications.
- Deploy and manage ML workloads across AWS, Azure or GCP.
- Implement model versioning, experiment tracking, model registries and reproducible ML environments.
- Containerise ML applications using Docker and deploy using Kubernetes or cloud-native services.
- Implement automated model testing, validation and deployment processes.
- Monitor model performance, data quality, drift, latency and infrastructure health in production.
- Work closely with Data Scientists, ML Engineers, Data Engineers, Software Engineers and Cloud Architects.
- Develop infrastructure-as-code using tools such as Terraform.
- Implement appropriate security, access control, secrets management and governance across ML platforms.
- Optimise ML infrastructure for performance, scalability and cost.
- Establish best practices around model governance, reproducibility and operational reliability.
- Support the adoption of Generative AI, LLM and Agentic AI solutions where relevant.
Technical Environment
Candidates should have experience with a combination of the following:
Cloud:
AWS, Azure or GCP
MLOps / ML Platforms:
MLflow, Kubeflow, SageMaker, Azure Machine Learning, Vertex AI, Databricks
DevOps:
Git, GitHub/GitLab, Jenkins, Azure DevOps, CI/CD
Containers & Infrastructure:
Docker, Kubernetes, Terraform
Data & ML:
Python, SQL, PySpark, Pandas, Scikit-learn, TensorFlow or PyTorch
Data Platforms:
Databricks, Snowflake, Data Lakes, Data Warehouses
Monitoring:
Prometheus, Grafana, CloudWatch, Azure Monitor or equivalent
AI:
Generative AI, LLMs, model serving, embeddings, vector databases and AI agents would be advantageous.
What We're Looking For
- 4+ years of experience in MLOps, ML Engineering, DevOps, Data Engineering or a closely related discipline.
- Strong Python and software engineering skills.
- Practical experience deploying machine learning models into production.
- Strong understanding of CI/CD, cloud infrastructure and automation.
- Experience working with containers and/or Kubernetes.
- Good understanding of data pipelines and modern data platforms.
- Experience with infrastructure-as-code, preferably Terraform.
- Strong understanding of model monitoring, versioning and governance.
- Comfortable working across technical teams and translating requirements into scalable solutions.
- A pragmatic, delivery-focused approach with strong problem-solving skills.
Advantageous
- Experience with LLMs, Generative AI or Agentic AI.
- Experience with Databricks, MLflow, Azure ML, AWS SageMaker or Google Vertex AI.
- Experience implementing enterprise AI platforms.
- Experience with real-time ML inference and model serving.
- Experience within highly regulated environments such as Financial Services, Insurance, Healthcare or Telecommunications.
The Opportunity
This is an opportunity to play a key role in building the infrastructure that enables production-grade AI and machine learning. You will work on meaningful technology initiatives, helping organisations move beyond AI experimentation and build secure, scalable and operational AI capabilities.