AI Jobs Map

Acendeo · Brazil

Sr. NVIDIA GPU Infrastructure Engineer / AI Install

seniorfull timePosted 18 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

embedded-systemsapache-airflowlinuxcudakubernetesrag

Description
Lead the hands-on bring-up of an enterprise, on-premises NVIDIA H200 (HGX 8-GPU SXM) environment built on HPE hardware with NVIDIA AI Enterprise. You will take this deployment from physical rack-and-cable installation through firmware validation and cluster initialization to hand off a production-ready environment to the client's AI engineering team.

English fluency is a MUST for this role! Only candidates with level C1 or C2 will be considered:

A1 Beginner

A2 Elementary

B1 Intermediate

B2 Upper-Intermediate

C1 Advanced

C2 Proficient

Responsibilities

- Validate facility readiness (power, cooling/airflow, rack weight, network drops) and execute physical racking and cabling of HPE H200 nodes.

- Flash and configure system firmware, BIOS/iLO, GPU VBIOS, NIC/HCA firmware, IOMMU, large-BAR, and performance profiles.

- Install and optimize the Linux OS, NVIDIA data-center driver, CUDA toolkit, Fabric Manager (NVSwitch), and MOFED GPUDirect/RDMA stack for InfiniBand/RoCE.

- Deploy container runtimes with NVIDIA Container Toolkit, Kubernetes, GPU Operator (and/or Base Command Manager), and configure Run:ai for workload scheduling and fractional-GPU sharing.

- Conduct burn-in testing and cluster validation, including NVLink topology checks, DCGM diagnostics, NCCL multi-GPU/multi-node benchmarks, and thermal/power soak tests.

- Support initial NVIDIA Inference Microservice (NIM) model-serving deployments and deliver a documented, repeatable environment with operational runbooks.

Requirements

- 5+ years of hands-on experience deploying NVIDIA data-center GPUs (Hopper H100/H200, HGX/SXM architectures).

- Deep technical expertise with NVLink, NVSwitch, Fabric Manager, and high-speed fabrics (InfiniBand or RoCE).

- Advanced Linux system administration skills (Ubuntu/RHEL), driver/kernel troubleshooting, and CUDA stack integration.

- Hands-on containerization and Kubernetes orchestration experience using GPU Operator.

- Proven ability to work independently in regulated enterprise environments with disciplined documentation habits.

Nice To Have

- Experience with HPE server platforms, NVIDIA AI Enterprise, Base Command Manager, or Run:ai.

- Familiarity with NIM inference serving engines and RAG pipeline architectures.

- Prior experience operating within telecom or regulated/sovereign-cloud environments.

- Willingness and ability to travel for extended on-site deployment rotations.