Our team members are the key to our company’s success, and their health and well-being, as well as that of their families, is very important to us. We offer a comprehensive benefits package that allows our team members stay healthy, plan for their future and maintain a healthy work-life balance. Benefits may vary with employment status. To see our fill list of Team Member Benefits please visit our career site: www.gotoworkhappy.com/benefits
Job Description:
We are looking for a highly skilled MLOps Engineer to support the end-to-end machine learning lifecycle, from experimentation to production deployment.
This role focuses on building scalable, reliable, and automated ML infrastructure, enabling data science teams to deliver production-ready models efficiently and confidently.
Key Responsibilities
-
Design, build, and maintain production-grade ML pipelines on Databricks
-
Operationalize ML models, including deployment, monitoring, and lifecycle management
-
Build and maintain CI/CD pipelines for ML workflows
-
Develop and manage real-time and streaming data pipelines
-
Collaborate closely with Data Scientists to productionize models efficiently
-
Implement model versioning, experiment tracking, and reproducibility
-
Define and enforce ML best practices, governance, and quality standards
-
Monitor model performance and data drift; implement automated retraining strategies
-
Optimize performance, scalability, and cost of distributed workloads
-
Contribute to platform design for low-latency inference and scalable serving
Required Qualifications (Must-Have)
-
Strong experience with Databricks (Workflows, MLflow, Delta Lake)
-
Deep expertise in Apache Spark (batch and streaming)
-
Advanced Python skills (production-quality code)
-
Hands-on experience with streaming / real-time systems
-
Proven experience designing and implementing CI/CD pipelines
Strong understanding of the ML lifecycle (training deployment monitoring
- retraining)
-
Experience building scalable, distributed data and ML pipelines
Nice-to-Have Skills
-
Experience with Snowflake
-
Knowledge of Kubernete
-
Experience with Docker
-
Familiarity with Terraform or other Infrastructure as Code tools
-
Experience with feature stores (e.g. Snowflake or Databricks Feature Store, etc.)
-
Experience with event-driven architectures (Kafka)
-
Experience with model serving frameworks and low-latency APIs
-
Monitoring and observability tools (ELK or similar)
-
Familiarity with A/B testing / experimentation frameworks
-
Experience with LLM deployment and serving
-
Knowledge of RBAC, security, and governance in data/ML platforms
-
Experience in cloud environments (Azure preferred)
What Success Looks Like
-
Fully automated, reliable ML pipelines from experimentation to production
-
High-quality, observable, and maintainable ML systems
-
Strong alignment between data science, engineering, and platform teams
-
Scalable infrastructure that supports both batch and real-time workloads
Example Use Cases You Will Support
-
Recommendation Systems (real-time / near real-time customer personalization)
-
LLM-based Products, including Text-to-SQL systems
-
Customer Personalization