About The Role
The Machine Learning Engineer will design, build, and operate production ML systems across the full model lifecycle, from data preparation and experimentation through deployment, monitoring, and continuous improvement. The work will span areas such as recommendation, ranking, forecasting, classification, NLP, and generative AI depending on product priorities.
Based in San Diego, CA with a remote work arrangement, the role partners with data scientists, software engineers, and platform teams to turn research prototypes into reliable services. Model quality, inference latency, scalability, and operational resilience are treated as equally important outcomes.
Key Responsibilities
- Design, train, and evaluate machine learning models using Python, PyTorch, TensorFlow, or scikit-learn for production use cases
- Build scalable data and feature pipelines with Python, SQL, Spark, and workflow orchestration tools such as Airflow or Kubeflow
- Deploy and serve models through AWS SageMaker, Kubernetes, or comparable cloud infrastructure, including model versioning, canary releases, and rollback procedures
- Develop reproducible training and experimentation workflows using MLflow, Weights & Biases, or equivalent tooling
- Monitor production models for latency, data drift, feature quality, bias, and performance regression with automated dashboards and alerts
- Optimize inference systems for throughput and cost using batching, caching, quantization, and appropriate serving frameworks
- Document technical decisions, write tested maintainable code, participate in architecture reviews, and mentor engineers on ML engineering practices
What We Are Looking For
- 3–8 years of experience in machine learning engineering, applied machine learning, or a closely related software engineering role, including production model deployment
- Strong Python skills and hands-on experience with at least one major ML framework, such as PyTorch, TensorFlow, or scikit-learn
- Proficiency in ML fundamentals including feature engineering, model selection, evaluation metrics, regularization, cross-validation, and error analysis
- Experience building data pipelines with SQL and Spark, along with a practical understanding of data quality, leakage prevention, and training-serving consistency
- Experience deploying ML systems on AWS, GCP, or Azure using containers, Kubernetes, CI/CD, and infrastructure or platform automation
- Bachelor’s or master’s degree in computer science, machine learning, statistics, mathematics, engineering, or a related technical field
- Bonus: Experience with LLM or generative AI systems, distributed training, real-time inference, feature stores, GPU optimization, or model observability platforms