Description
Lead the hands-on bring-up of an enterprise, on-premises NVIDIA H200 (HGX 8-GPU SXM) environment built on HPE hardware with NVIDIA AI Enterprise. You will take this deployment from physical rack-and-cable installation through firmware validation and cluster initialization to hand off a production-ready environment to the client's AI engineering team.
English fluency is a MUST for this role! Only candidates with level C1 or C2 will be considered:
A1 Beginner
A2 Elementary
B1 Intermediate
B2 Upper-Intermediate
C1 Advanced
C2 Proficient
Responsibilities
- Validate facility readiness (power, cooling/airflow, rack weight, network drops) and execute physical racking and cabling of HPE H200 nodes.
- Flash and configure system firmware, BIOS/iLO, GPU VBIOS, NIC/HCA firmware, IOMMU, large-BAR, and performance profiles.
- Install and optimize the Linux OS, NVIDIA data-center driver, CUDA toolkit, Fabric Manager (NVSwitch), and MOFED GPUDirect/RDMA stack for InfiniBand/RoCE.
- Deploy container runtimes with NVIDIA Container Toolkit, Kubernetes, GPU Operator (and/or Base Command Manager), and configure Run:ai for workload scheduling and fractional-GPU sharing.
- Conduct burn-in testing and cluster validation, including NVLink topology checks, DCGM diagnostics, NCCL multi-GPU/multi-node benchmarks, and thermal/power soak tests.
- Support initial NVIDIA Inference Microservice (NIM) model-serving deployments and deliver a documented, repeatable environment with operational runbooks.
Requirements
- 5+ years of hands-on experience deploying NVIDIA data-center GPUs (Hopper H100/H200, HGX/SXM architectures).
- Deep technical expertise with NVLink, NVSwitch, Fabric Manager, and high-speed fabrics (InfiniBand or RoCE).
- Advanced Linux system administration skills (Ubuntu/RHEL), driver/kernel troubleshooting, and CUDA stack integration.
- Hands-on containerization and Kubernetes orchestration experience using GPU Operator.
- Proven ability to work independently in regulated enterprise environments with disciplined documentation habits.
Nice To Have
- Experience with HPE server platforms, NVIDIA AI Enterprise, Base Command Manager, or Run:ai.
- Familiarity with NIM inference serving engines and RAG pipeline architectures.
- Prior experience operating within telecom or regulated/sovereign-cloud environments.
- Willingness and ability to travel for extended on-site deployment rotations.