NVIDIA develops accelerated computing platforms that support large language model innovation and deployment. The company is seeking a Senior DL Performance Efficiency Architect to lead cross-layer efforts that improve LLM efficiency, guide model-system-hardware co-design, establish an efficiency roadmap, and drive optimization projects through production.
Responsibilities
- Lead cross-layer efforts to improve the efficiency of large language models across model architecture, training and inference systems
- Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure, and identify opportunities for model-system-hardware co-design
- Establish a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment
- Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence future model, software, and hardware roadmaps