Job Description:
Key Responsibilities:
1. Platform Architecture & Ownership - Design and own end-to-end ML platform architecture (data ? training ? deployment ? monitoring) - Define and enforce best practices for scalable and secure ML systems - Standardize MLOps + DevOps frameworks and processes
2. Model Deployment & Serving - Deploy and manage ML/LLM models on GPU-based on-prem infrastructure - Optimize inference performance (latency, throughput, batching) - Implement model versioning, A/B testing, and rollback strategies
3. CI/CD & Automation - Design and implement CI/CD pipelines for ML models, APIs, and data workflows - Enable automated testing, deployment, and release management
4. Infrastructure & Containerization - Manage Linux-based (RHEL preferred) on-prem infrastructure - Containerize applications using Docker - Deploy and orchestrate workloads using Kubernetes / OpenShift - Operate within restricted or air-gapped environments
5. Data & System Integration - Build pipelines integrating structured databases and high-volume logs/streaming data - Support batch and real-time inference architectures
6. Monitoring, Observability & Reliability - Implement end-to-end observability (model + infra) - Use tools like Prometheus, Grafana, ELK stack - Ensure high availability, SLA adherence, and incident response
7. GenAI & Advanced ML Systems - Deploy RAG pipelines and vector databases - Manage LLM serving frameworks - Work with agent orchestration frameworks
8. Leadership & Collaboration - Mentor engineers on MLOps and DevOps best practices - Collaborate with cross-functional teams - Drive design reviews and production readiness
Required Skills: -
Strong Python and scripting (Bash)
Deep understanding of ML lifecycle and productionization - Experience deploying ML/LLM systems in production
Linux, Docker, Kubernetes/OpenShift
CI/CD tools (Jenkins/GitLab CI) - SQL and data pipeline experience
Good to Have: -
GPU optimization knowledge - MLflow / Kubeflow - Terraform / Ansible
Experience in on-prem or restricted environments
Experience: - 6+ years in MLOps / DevOps / Platform Engineering
Proven experience scaling production ML systems
Ideal Candidate: A hands-on platform architect who can operate across ML systems and infrastructure, driving automation, scalability, and reliability.
Pay: ₹50,000.00 - ₹110,000.00 per month
Work Location: In person