AI Jobs Map

UMA · Paris, Île-de-France, France

Compute Infrastructure Lead

directorfull timePosted 3 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

observabilityprometheusgrafanamlflowpytorchlinuxkubernetespythonllmray

Your Mission

As Compute Infrastructure Lead, you will own and scale the compute backbone of UMA: the systems that provision, schedule, and run training, evaluation, and data-processing workloads — reliably, efficiently, and at scale — so our models can go from research to production without the cluster becoming the bottleneck.

This is a hands-on, high-impact role. We already train on a dedicated GPU cluster with a working training stack and a strong team behind it, so you won't be starting from zero — but you'll have the mandate to shape the architecture that takes us from a research cluster to a production-scale, multi-provider fleet, and to production-grade reliability as we start deploying POCs with industry partners. You'll take ownership of the compute platform end to end — multi-provider capacity, scheduling, distributed training / eval / processing, virtualized developer environments, observability, cost, and the tooling researchers and engineers actually use — and, if that's where you want to go, grow into leading the compute infrastructure team as it scales.

The technical problem is unusually rich for this stage. We are de-risking a stack built on pre-training and online RL, then industrializing it: heterogeneous hardware (training GPUs, cheaper eval and processing GPUs, CPU), a real-time learning loop, orchestrated data processing, interactive VMs on the cluster, and a fleet that will grow fast. Much of what we need — a true multi-provider compute fabric with elastic scheduling, dynamic checkpointing, unified observability, and jobs that resume themselves — does not exist off the shelf. Data infrastructure (datasets, storage, versioning) is owned by a sister role; this role is compute, including how processing jobs actually run on it.

Key responsibilities :

- Own our compute platform end to end — from provisioning GPU capacity across cloud providers to keeping training, eval, and processing jobs running at high utilization, with reliability, cost, and researcher velocity as first-class goals

- Build a multi-provider management layer so we can place, burst, and fail over workloads across GPU clouds, hyperscalers, and HPC without rewriting jobs

- Design and operate the cloud scheduler — quotas, priority, preemption, topology-aware placement, and dynamic checkpointing so jobs survive node failure, preemption, and provider switches

- Stand up a distributed compute framework for training and evaluation on heterogeneous hardware (e.g. Ray / similar), including the real-time / online-learning path

- Orchestrate data-processing workloads at scale — CPU and cheaper GPUs, batch and streaming — so post-processing, dataset jobs, and training share one reliable compute fabric instead of ad-hoc scripts

- Deliver virtualized GPU/CPU dev sessions (VMs on the cluster) so engineers iterate interactively on the same hardware and software they train on, without burning dedicated boxes

- Build observability that works from any provider — system metrics (Prometheus, Grafana), job traces and logs, and model metrics (e.g. MLflow) — so a hung NCCL job, a silent GPU, or a broken training curve is diagnosable in minutes, not days

- Own capacity, cost, and provider relationships as a technical lead: forecast demand, pick the right mix of hardware and contracts, and help negotiate pricing and terms. This is not a commercial role — but procurement is part of making the infra succeed

- Help set production-grade practices (testing, reliability, fast iteration) as we move from R&D to partner POCs, and grow into leading the compute infra team if that's the path you want

What You Bring To The Table

- 8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering, at a senior, lead, or staff level

- Proven track record building and operating infrastructure for large-scale AI model training — not inference-only. Multi-node GPU clusters, distributed training (PyTorch / NCCL or equivalent), and keeping long-running jobs healthy at scale

- Deep, hands-on experience with GPU clouds and cluster operations: provisioning, Linux, high-performance networking (InfiniBand / RoCE), storage for training, utilization, and GPU/node failure modes

- Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot, or similar) — including checkpointing, elasticity, and preemption — so GPUs stay busy and jobs come back from failure

- Treat observability, reliability, and cost as core engineering concerns, not afterthoughts

- Strong Python and systems engineering, with the taste to build tooling that researchers actually want to use

- Experience working with GPU providers on capacity and commercial terms — you can read a contract, push on price and availability, and still be the person who debugs the cluster at 2am

- Ability to reason about systems end-to-end — performance, scalability, reliability, cost — and make and defend the right trade-offs

- Thrive in a hands-on, fast-paced startup, building from a real (but small) cluster toward a production fleet: autonomous, rigorous, execution-driven, easy to work with, and broadly curious about AI and systems

- Bonus : online / continuous RL, real-time training loops, or other always-on learning systems

- Bonus : multi-cloud / multi-provider fabrics (SkyPilot, similar), Ray / Anyscale, HPC centers, VM-based GPU workstations / interactive cluster sessions, or standing up clusters from tens to hundreds of nodes

- Bonus : robotics, autonomous vehicles, or other embodied/physical-AI training stacks — adjacent large-scale training (LLMs, multimodal, AV) counts strongly; robotics itself is not required

- Bonus : public projects, open-source contributions, maintained tools, or technical writing

- We value exceptional builders over perfect resumes. If you have a world-class track record building training infrastructure at scale and the drive to build the compute backbone that lets a robotics company scale, we strongly encourage you to apply — even if you don't tick every box. Robotics experience is a plus, not a requirement.

More jobs at UMA