Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing. An era in which GPUs serve as the brains of computers, robots, and self-driving cars that can understand and interact with the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.
NVIDIA is hiring a Senior Software Architect in the Agent Harness & Runtime Engineering team to build foundational systems for the next generation of agentic AI. Our team works at the intersection of AI and systems engineering, tackling challenges across agent runtimes, inference, evaluation, and large-scale execution. We are looking for someone with the AI breadth to quickly understand emerging problems, the systems depth to architect solutions at scale, and the engineering strength to build them. This is an opportunity to work across the agentic AI stack and translate rapidly evolving AI capabilities into scalable, production-quality systems.
What You'll Be Doing
- Architect and build scalable, reliable systems for agentic AI, spanning agent runtimes, harnesses, inference, evaluation, and orchestration.
- Solve large-scale systems challenges across distributed execution, data/ETL pipelines, HPC, cloud, Kubernetes, and GPU compute environments.
- Optimize systems for reliability, scalability, performance, resource utilization, and developer experience.
- Prototype emerging ideas, write high-quality production code, and provide technical leadership across research, engineering, product, and infrastructure teams.
What We Need To See
- BS, MS, or equivalent experience in Computer Science, Computer Engineering, AI, or a related field, with 12+ years of relevant industry experience.
- Strong foundation in modern AI, including LLMs, multimodal models, inference, agentic AI, and evaluation.
- Deep expertise in software architecture, distributed systems, and large-scale systems design.
- Strong hands-on programming skills in Python, C++, Go, Rust, or similar languages, with a track record of building high-quality production software.
- Background building large-scale systems involving distributed execution, data processing/ETL, workflow orchestration, HPC, cloud, or GPU infrastructure.
- Ability to reason across the stack—from AI model behavior and application logic to runtime, compute, and infrastructure.
- Proven ability to take ambiguous, complex technical problems from architecture through implementation and production.
Ways To Stand Out From The Crowd
- Knowledge of LLM/VLM inference, model serving, model routing, or inference optimization.
- Background in AI evaluation, benchmarking, experimentation, or large-scale AI infrastructure.
- Track record of building reusable platforms supporting heterogeneous AI workloads and compute environments.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 13, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.