AI Jobs Map

Blue Signal Search · United States

Head of GPU Cloud

seniorfull timePosted 2 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

system-designvllmllmkubernetesobservabilitytime-series

Head of GPU Cloud

Remote - Nationwide

Our client is building the software foundation for a next generation accelerated computing platform designed to make large-scale AI infrastructure easier to consume, operate, and scale. They are seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming significant GPU capacity into high-performance cloud services for AI workloads. This is an opportunity to shape platform architecture, production inference, developer experiences, and engineering strategy at a stage where technical decisions will directly influence customer adoption, infrastructure economics, and long-term growth.

This Role Offers

- Opportunity to define the architecture and operating model behind large scale AI inference services.

- Direct influence over GPU utilization, platform economics, customer experience, and technical strategy.

- Close collaboration with leaders across infrastructure, networking, product, operations, and commercial functions.

- A highly technical environment where software engineering intersects with accelerated computing, distributed systems, and AI infrastructure.

Focus

- Build and scale the engineering organization behind a nationwide GPU cloud platform, with ownership across inference services, orchestration, APIs, platform reliability, and technical execution.

- Set the architecture for moving AI workloads efficiently from customer request to accelerator, including routing, placement, model lifecycle management, caching, and memory aware scheduling.

- Lead production model serving and optimization across technologies such as vLLM, TensorRT LLM, and TGI, improving throughput, latency, availability, accelerator utilization, and cost.

- Own Kubernetes based GPU infrastructure strategy across workload scheduling, elasticity, observability, deployment automation, multi-tenant isolation, and operational resilience.

- Develop platform capabilities for hosted models, private inference environments, customized deployments, usage measurement, customer controls, and developer facing services.

- Design orchestration policies that match workloads to accelerators based on memory requirements, performance objectives, capacity, cluster topology, and infrastructure economics.

- Partner with networking, systems, data center, product, and commercial leaders to align software decisions with accelerator architecture, fabric performance, storage, and customer requirements.

- Establish engineering standards, service objectives, capacity planning, incident readiness, and team accountability while recruiting and developing senior technical talent.

Skill Set

- 12 or more years of progressive software engineering experience, including substantial leadership responsibility across cloud platforms, distributed infrastructure, HPC, or similarly complex production systems.

- 5 or more years leading engineering teams responsible for business critical infrastructure, platform services, or other mission critical technology products.

- Demonstrated production experience operating large scale model inference using vLLM, TensorRT LLM, TGI, or equivalent serving stacks.

- Strong expertise in model serving optimization, including dynamic batching, decoding acceleration, reduced precision execution, compilation, memory reuse, and request scheduling.

- Advanced knowledge of Kubernetes and containerized infrastructure, including scheduling, elasticity, telemetry, deployment practices, and production reliability.

- Practical accelerator infrastructure knowledge covering GPU memory behavior, high bandwidth networking, InfiniBand, RoCE, storage performance, and cluster topology.

- Experience architecting highly available distributed platforms with automated resource allocation, programmatic interfaces, multi customer support, and detailed consumption measurement.

- Strong technical and leadership judgment with the ability to balance performance, reliability, security, customer experience, and infrastructure economics across multidisciplinary teams.

Additional Experience That Stands Out

- Leadership experience in GPU cloud, AI infrastructure, hosted model platforms, or accelerated computing environments.

- Experience delivering elastic inference services, dedicated AI capacity, model customization workflows, or managed AI products.

- Familiarity with open model ecosystems and the operational differences among model families, serving configurations, and hardware profiles.

- Experience creating consistent developer interfaces across multiple model backends.

- Track record improving accelerator utilization, workload density, capacity forecasting, and compute economics across multiple GPU generations.

- Experience scaling engineering organizations in fast moving environments where software requirements and infrastructure capacity evolve together.

About Blue Signal:

Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS

More jobs at Blue Signal Search