AI Jobs Map

Rafay Systems · United States

Principal Solutions Architect

Hybriddirectorfull timePosted yesterday
Apply on IndeedIndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

artificial-intelligencemlopsdata-sciencekubernetesetlllmvllmobservabilityopentelemetrypythonbashawsazuregcpprometheusgrafanapytorchtensorflow

About the Role

Rafay seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds in deploying, operating, and scaling AI/ML workloads on our GPU Platform-as-a-Service offering. This is a hybrid technical leadership and people-management role: alongside hands-on, customer-facing architecture work, this person will build, lead, and grow a team of Solutions Architects, collaborating with platform engineering, MLOps, data science, and infrastructure teams to architect production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments.

Key Responsibilities

Team Leadership & People Management

-
Recruit, hire, and onboard Solutions Architects as the team scales

-
Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support

-
Set individual and team goals; conduct regular 1:1s and performance reviews

-
Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning

-
Foster an inclusive, high-performing team culture aligned with Rafay's values

-
Manage team capacity, prioritization, and staffing against customer and project demand

-
Partner with sales, engineering, and executive leadership on hiring plans and team structure

-
Mentor and upskill both direct reports and junior team members across the broader organization

Technical & Customer-Facing Responsibilities

-
Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines

-
Develop reference architectures for GPU cluster deployment and LLM serving infrastructure

-
Evaluate inference serving frameworks including vLLM, TGI, and Triton

-
Advise on GPU fabric topology options for distributed training scenarios

-
Design observability strategies using DCGM, OpenTelemetry, and eBPF

-
Translate infrastructure requirements into actionable platform designs

-
Deliver technical presentations, workshops, and proof-of-concept engagements

-
Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling

-
Partner with customer stakeholders to understand workload requirements

-
Architect networking, identity management, observability, and security integrations

-
Monitor and troubleshoot production environments for GPU utilization and cluster health

-
Lead root cause analysis for complex customer issues

-
Document reference architectures and implementation best practices

Required Qualifications

-
8+ years in infrastructure, platform, or solutions engineering roles

-
3+ years focused on AI/ML infrastructure or MLOps

-
2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff

-
Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture

-
Deep Kubernetes expertise including cluster lifecycle and RBAC

-
Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)

-
Proficiency with distributed training concepts (NCCL, tensor parallelism)

-
Experience with LLM inference serving and optimization

-
Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics

-
Strong scripting and automation skills (Python, Bash, Go preferred)

-
Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders

-
Experience with AWS, Azure, or GCP platforms

-
Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry

-
Understanding of GPU-based workloads and model serving

-
Proven troubleshooting capabilities for infrastructure issues

-
Excellent communication, coaching, and customer-facing skills

Preferred Qualifications

-
Experience building a Solutions Architecture or technical pre-sales team from the ground up

-
Formal people-management training or leadership certification

-
Enterprise customer support experience in cloud-native environments

-
Familiarity with PyTorch and TensorFlow frameworks

-
Experience with Run:AI and Slurm

-
GPU scheduling and autoscaling expertise

-
Multi-tenant Kubernetes environment knowledge

-
MLOps platform experience

-
Technical workshop leadership experience

-
Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect)

-
Understanding of multi-tenant GPU isolation technologies

Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

More jobs at Rafay Systems