Build and operate scalable AI/ML platforms, CI/CD pipelines and cloud infrastructure for ML and LLM workloads, with a focus on automation, security, observability and model lifecycle management for our client.
Key Responsibilities
- Design and manage CI/CD pipelines for ML/LLM workloads using Jenkins, GitHub Actions and Kubeflow.
- Build reusable deployment templates, workflows and platform components.
- Automate infrastructure using Terraform, Helm and Kubernetes.
- Deploy ML/LLM applications on Kubernetes and EKS/AKS/GKE.
- Implement monitoring and observability using ELK, Prometheus and Grafana.
- Ensure security, compliance and cost governance across cloud environments.
- Collaborate with Data Scientists, ML Engineers and DevOps teams.
- Mentor engineers and continuously improve platform reliability and scalability.
Technical Skills
- 8+ years total experience with 5+ years in MLOps/LLMOps or related areas.
- Strong CI/CD expertise: Jenkins, GitHub Actions.
- Python, Shell, Docker/Dockerfiles and Terraform.
- Kubernetes and managed Kubernetes services.
- Terraform, Helm and Infrastructure as Code.
- Kubeflow for ML workflow orchestration.
- ELK Stack, Prometheus and Grafana.
- Strong AWS and Azure experience.
Understanding of ML/LLM deployment and model lifecycle management