About the Role
Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads.
Key Responsibilities
Core Implementation Responsibilities
-
Serve as primary technical lead for customer implementation engagements
-
Partner with customers on requirements gathering, architecture reviews, and solution design
-
Design and configure Rafay platform capabilities across public, private, and hybrid cloud environments
-
Troubleshoot complex infrastructure, networking, Kubernetes, and virtualization issues
-
Collaborate with Customer Success, Product, and Engineering teams on technical challenges
-
Manage customer issues within SLAs while maintaining expectations
-
Reproduce and analyze customer-reported issues; communicate findings to internal teams
-
Develop technical documentation, implementation guides, and runbooks
-
Mentor junior engineers and contribute to process improvements
-
Stay current on Rafay platform capabilities and emerging industry trends
GPU & AI/ML Infrastructure
-
Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup
-
Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations
-
Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2)
-
Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling
-
Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200)
Minimum Qualifications
-
10+ years of experience in customer-facing technical roles, including implementation or consulting
-
Hands-on AI/ML engineering experience with model inference and deployment workflows
-
Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
-
Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals
-
Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton
-
Strong Kubernetes and cloud-native environment expertise
-
Excellent written and verbal communication abilities
-
Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms
-
Infrastructure automation and IaC (Terraform preferred) experience desired
-
Linux systems administration and distributed systems knowledge
-
Proven ability to lead technical projects and manage multiple engagements
-
Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience
Preferred Qualifications
-
Experience with Run:AI and Slurm for GPU scheduling
-
GPU autoscaling and multi-tenant GPU isolation expertise
-
Familiarity with DCGM-based GPU observability and Prometheus/Grafana dashboards
-
Relevant certifications (CKA, CKAD, or NVIDIA certifications)
Why Join Rafay?
Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.