Sr. Software Development Engineer (ML Infrastructure / MLOps)
Location: Remote
Contract: 12+ Months
We are seeking a Sr. Software Development Engineer with strong experience in ML Infrastructure, MLOps, and large-scale distributed systems. This role will focus on building and optimizing scalable ML platforms, training pipelines, serving infrastructure, and automation systems supporting production AI/ML workloads serving millions of users.
Required Skills
- 5+ years of experience in ML Infrastructure, MLOps, or ML Platform Engineering
- Strong programming skills in Python and Java
- Hands-on experience with PyTorch, TensorFlow
- Expertise with Kubernetes (GKE), Docker, Terraform, and CI/CD
- Experience with GCP (AWS experience also considered)
- Building and supporting large-scale ML training and deployment pipelines
- Experience with distributed systems, workflow automation, monitoring, and scalability optimization
- Proven production-level experience supporting high-volume ML environments
Preferred Skills
- LLM/AI infrastructure and serving (vLLM, Triton, TensorRT)
- Ray, DeepSpeed, Vertex AI, Weights & Biases
- GPU optimization, distributed training, quantization, pruning
- Agentic AI, Gemini, ChatGPT, prompt engineering
- Data labeling, feature generation, synthetic data, and ML monitoring systems
Ideal Candidate
- Strong background in ML Infrastructure Engineer or MLOps Engineer roles
- Experience improving system performance, GPU utilization, scalability, and cost optimization
- Hands-on ownership of production ML platforms and large-scale distributed systems
- Not suitable for pure research, academic ML, or theoretical Data Science profiles