About The Role
The MLOps Engineer will build and operate the infrastructure that moves machine learning and generative AI systems from experimentation into reliable production services. The role covers training pipelines, model registries, deployment automation, inference environments, observability, and governance across cloud-based platforms.
This engineer will partner with ML engineers, data scientists, and platform teams to improve deployment speed, model reliability, and operational efficiency. The work is highly hands-on, with direct ownership of Kubernetes workloads, CI/CD pipelines, infrastructure as code, and monitoring for model and system performance.
Key Responsibilities
- Build and maintain automated ML pipelines for data validation, training, evaluation, model registration, and production deployment using tools such as Kubeflow, MLflow, Airflow, or equivalent
- Deploy and operate real-time and batch inference services on AWS, GCP, or Azure using Docker, Kubernetes, Helm, and infrastructure-as-code tools such as Terraform
- Develop CI/CD workflows for ML and GenAI applications, including automated testing, container builds, model promotion, rollback, and environment management
- Implement observability for model and platform health, including latency, throughput, resource utilization, data drift, model quality, and inference cost
- Collaborate with ML engineers and data scientists to standardize reproducible training environments, feature pipelines, experiment tracking, and model versioning
- Improve reliability and efficiency through capacity planning, GPU utilization, autoscaling, caching, performance tuning, and incident response
- Establish security and governance controls for model artifacts, datasets, secrets, access permissions, and auditability across the ML platform
What We Are Looking For
- 3–8 years of experience in MLOps, machine learning engineering, platform engineering, or a closely related discipline, including hands-on ownership of production ML systems
- Strong Python and Linux skills, with experience writing maintainable automation, services, and deployment tooling
- Production experience with Docker, Kubernetes, Helm, and at least one major cloud platform such as AWS, GCP, or Azure
- Proficiency with infrastructure as code and CI/CD tools, including Terraform, GitHub Actions, GitLab CI, Jenkins, or comparable technologies
- Experience with ML lifecycle and orchestration platforms such as MLflow, Kubeflow, Airflow, Argo Workflows, SageMaker, Vertex AI, or Azure ML
- Working knowledge of model serving, REST or gRPC APIs, batch inference, feature stores, data validation, and monitoring for model drift and performance degradation
- Bachelor’s degree in computer science, engineering, mathematics, or a related technical field; equivalent practical experience is also considered. Bonus: experience supporting LLM inference, GPU workloads, Ray, vLLM, vector databases, or production RAG systems