Head of GPU Cloud
Remote - Nationwide
Our client is building the software foundation for a next generation accelerated computing platform designed to make large-scale AI infrastructure easier to consume, operate, and scale. They are seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming significant GPU capacity into high-performance cloud services for AI workloads. This is an opportunity to shape platform architecture, production inference, developer experiences, and engineering strategy at a stage where technical decisions will directly influence customer adoption, infrastructure economics, and long-term growth.
This Role Offers
- Opportunity to define the architecture and operating model behind large scale AI inference services.
- Direct influence over GPU utilization, platform economics, customer experience, and technical strategy.
- Close collaboration with leaders across infrastructure, networking, product, operations, and commercial functions.
- A highly technical environment where software engineering intersects with accelerated computing, distributed systems, and AI infrastructure.
Focus
- Build and scale the engineering organization behind a nationwide GPU cloud platform, with ownership across inference services, orchestration, APIs, platform reliability, and technical execution.
- Set the architecture for moving AI workloads efficiently from customer request to accelerator, including routing, placement, model lifecycle management, caching, and memory aware scheduling.
- Lead production model serving and optimization across technologies such as vLLM, TensorRT LLM, and TGI, improving throughput, latency, availability, accelerator utilization, and cost.
- Own Kubernetes based GPU infrastructure strategy across workload scheduling, elasticity, observability, deployment automation, multi-tenant isolation, and operational resilience.
- Develop platform capabilities for hosted models, private inference environments, customized deployments, usage measurement, customer controls, and developer facing services.
- Design orchestration policies that match workloads to accelerators based on memory requirements, performance objectives, capacity, cluster topology, and infrastructure economics.
- Partner with networking, systems, data center, product, and commercial leaders to align software decisions with accelerator architecture, fabric performance, storage, and customer requirements.
- Establish engineering standards, service objectives, capacity planning, incident readiness, and team accountability while recruiting and developing senior technical talent.
Skill Set
- 12 or more years of progressive software engineering experience, including substantial leadership responsibility across cloud platforms, distributed infrastructure, HPC, or similarly complex production systems.
- 5 or more years leading engineering teams responsible for business critical infrastructure, platform services, or other mission critical technology products.
- Demonstrated production experience operating large scale model inference using vLLM, TensorRT LLM, TGI, or equivalent serving stacks.
- Strong expertise in model serving optimization, including dynamic batching, decoding acceleration, reduced precision execution, compilation, memory reuse, and request scheduling.
- Advanced knowledge of Kubernetes and containerized infrastructure, including scheduling, elasticity, telemetry, deployment practices, and production reliability.
- Practical accelerator infrastructure knowledge covering GPU memory behavior, high bandwidth networking, InfiniBand, RoCE, storage performance, and cluster topology.
- Experience architecting highly available distributed platforms with automated resource allocation, programmatic interfaces, multi customer support, and detailed consumption measurement.
- Strong technical and leadership judgment with the ability to balance performance, reliability, security, customer experience, and infrastructure economics across multidisciplinary teams.
Additional Experience That Stands Out
- Leadership experience in GPU cloud, AI infrastructure, hosted model platforms, or accelerated computing environments.
- Experience delivering elastic inference services, dedicated AI capacity, model customization workflows, or managed AI products.
- Familiarity with open model ecosystems and the operational differences among model families, serving configurations, and hardware profiles.
- Experience creating consistent developer interfaces across multiple model backends.
- Track record improving accelerator utilization, workload density, capacity forecasting, and compute economics across multiple GPU generations.
- Experience scaling engineering organizations in fast moving environments where software requirements and infrastructure capacity evolve together.
About Blue Signal:
Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS