About The Role
The MLOps Engineer will build and operate the infrastructure that moves machine learning models from experimentation into reliable production systems. The role covers training and inference pipelines, model registries, deployment automation, observability, and cloud-native platform engineering across AWS or comparable environments.
Working alongside ML engineers, data scientists, and platform engineers, this role will improve deployment velocity without compromising reliability, security, or model quality. The team is focused on repeatable workflows for batch and real-time inference, with clear monitoring for latency, cost, data drift, and performance regression.
Key Responsibilities
- Build and maintain automated CI/CD pipelines for model training, validation, packaging, and deployment using GitHub Actions, GitLab CI, or Jenkins
- Deploy and scale batch and real-time inference services on Kubernetes using Docker, Helm, and infrastructure-as-code tools such as Terraform
- Implement ML workflows with platforms and frameworks such as Kubeflow, MLflow, Airflow, or AWS SageMaker
- Create model observability systems covering service health, latency, throughput, resource utilization, data drift, and prediction quality
- Manage model and dataset versioning, artifact storage, approval workflows, rollback procedures, and reproducible release processes
- Improve platform reliability through automated testing, security controls, incident response, capacity planning, and cost optimization
- Partner with ML and data engineering teams to standardize production interfaces, feature pipelines, environment configuration, and operational runbooks
What We Are Looking For
- 3–8 years of experience in MLOps, ML platform engineering, DevOps, or backend infrastructure, including experience supporting machine learning workloads in production
- Strong Python programming skills and experience building production services, automation, or data and ML pipelines
- Hands-on experience with Kubernetes, Docker, Helm, and CI/CD systems; familiarity with Terraform or another infrastructure-as-code framework
- Experience deploying workloads on AWS, GCP, or Azure, including services for compute, storage, networking, secrets, and monitoring
- Practical knowledge of ML lifecycle tooling such as MLflow, Kubeflow, SageMaker, Vertex AI, or equivalent platforms
- Understanding of model serving patterns, batch and online inference, feature stores, experiment tracking, monitoring, and data or model drift
- Bachelor’s degree in computer science, engineering, mathematics, or a related technical field; equivalent professional experience is acceptable. Bonus: experience with GPU orchestration, Ray, KServe, Argo Workflows, Prometheus, Grafana, and SRE practices