AI Jobs Map

D24 Search · Boston, MA

Machine Learning Engineer

seniorfull time$250,000 – $350,000 / yearPosted 20 days ago
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

vllmobservabilitypytorchpythoncudagcpawsterraformmachine-learningdata-structures

ML Engineer - Inference

Boston

$250K - $350K + equity

My client is a well-funded stealth AI-for-physics startup with roots at Harvard, MIT, Johns Hopkins, Oxford, the Institute for Advanced Study, and the Perimeter Institute, building AI systems that can discover new physics at scale. The company sits at the intersection of frontier AI research and fundamental science, combining cutting-edge models with deep physics expertise to unlock new scientific knowledge and real-world impact.

About the Role

The Member of Technical Staff (ML Engineer) builds and runs the training and inference systems that turn our research into things that work at scale, and that the rest of Engineering can build on. Hands-on: writing the code, not just the design doc, and the first call when a training job stalls or an inference path breaks.

Key Responsibilities

- Own the training and inference infrastructure that we depends on: distributed training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, or comparable) for both proprietary models and self-hosted inference.

- Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher's time goes into the science instead of the plumbing.

- Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else.

- Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are yours to close, not someone else's ticket.

Requirements

- 3+ years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.

- Hands-on with distributed training (multi-GPU or multi-node, using PyTorch, Ray, or comparable) and model-serving systems (vLLM, SGLang, Triton, or comparable).

- Strong software engineering fundamentals - can build a service that other engineers and researchers depend on daily, not a script that worked once.

- Enough ML fluency to work productively with AI researchers - understands training loops, reward signals, and inference-time behavior well enough to debug them, even without designing the algorithms.

- Strong programming in Python; comfortable operating production systems.

Bonus Skills

- Experience building internal platform tools such as training-as-a-service APIs, inference gateways, or job schedulers.

- Background in GPU infrastructure, CUDA, or performance engineering for ML workloads.

- Experience with cloud infrastructure (GCP, AWS) and infrastructure as code (Terraform or comparable).

- Prior work embedded alongside a research team, turning research code into production systems

More jobs at D24 Search

Similar roles in Boston