Position: ML Ops Engineer
Location: Concord, CA
Duration: 12 Months
Job Summary
We are seeking an experienced ML Ops Engineer to design, build, and support scalable, secure, and production-ready machine learning platforms across cloud and on-premises environments. The ideal candidate will have strong expertise in MLOps, Kubernetes, cloud platforms, automation, and reliability engineering
Required Skills:
- 8+ years in MLOps, Platform Engineering, DevOps, or SRE
- Strong GCP (Google Cloud Platform) experience
- Kubernetes (GKE/OpenShift)
- Python automation and platform development
- MLOps platforms and ML lifecycle management
- CI/CD pipelines and infrastructure automation
- Monitoring, observability, logging, and incident management
- Security, compliance, and reliability engineering
Preferred Skills:
- Enterprise-scale ML platform architecture
- AWS/Azure multi-cloud experience
- GenAI/AI workload deployment and support
- Terraform, Ansible, Infrastructure as Code
- Model serving, feature stores, model monitoring
- Technical leadership and mentoring