AI Jobs Map

Rafay Systems · United States

Senior Solutions Engineer

Hybridseniorfull timePosted yesterday
Apply on IndeedIndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

kubernetesartificial-intelligencellmvllmobservabilityopentelemetryterraformlinuxsystem-designprometheusgrafana

About the Role

Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads.

Key Responsibilities

Core Implementation Responsibilities

-
Serve as primary technical lead for customer implementation engagements

-
Partner with customers on requirements gathering, architecture reviews, and solution design

-
Design and configure Rafay platform capabilities across public, private, and hybrid cloud environments

-
Troubleshoot complex infrastructure, networking, Kubernetes, and virtualization issues

-
Collaborate with Customer Success, Product, and Engineering teams on technical challenges

-
Manage customer issues within SLAs while maintaining expectations

-
Reproduce and analyze customer-reported issues; communicate findings to internal teams

-
Develop technical documentation, implementation guides, and runbooks

-
Mentor junior engineers and contribute to process improvements

-
Stay current on Rafay platform capabilities and emerging industry trends

GPU & AI/ML Infrastructure

-
Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup

-
Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations

-
Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2)

-
Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling

-
Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200)

Minimum Qualifications

-
10+ years of experience in customer-facing technical roles, including implementation or consulting

-
Hands-on AI/ML engineering experience with model inference and deployment workflows

-
Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)

-
Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals

-
Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton

-
Strong Kubernetes and cloud-native environment expertise

-
Excellent written and verbal communication abilities

-
Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms

-
Infrastructure automation and IaC (Terraform preferred) experience desired

-
Linux systems administration and distributed systems knowledge

-
Proven ability to lead technical projects and manage multiple engagements

-
Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience

Preferred Qualifications

-
Experience with Run:AI and Slurm for GPU scheduling

-
GPU autoscaling and multi-tenant GPU isolation expertise

-
Familiarity with DCGM-based GPU observability and Prometheus/Grafana dashboards

-
Relevant certifications (CKA, CKAD, or NVIDIA certifications)

Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

More jobs at Rafay Systems