Director | AI / ML | Delhi | Engineering | Hybrid Cloud Engineering
- Job requisition ID : 111374
- Location: Delhi
- Entity: Deloitte Touche Tohmatsu India LLP
Infrastructure Management: Build and maintain on-premise hardware clusters (NVIDIA DGX, GPU nodes) and storage.
GPU Orchestration: Configure GPU scheduling, resource quotas, and multi-instance GPU (MIG) slices using tools like Run:ai or Kubernetes.
Model Serving Pipelines: Deploy and optimize inference backends such as vLLM, Triton, or TensorRT-LLM.
Platform Operations: Implement containerization, networking, Role-Based Access Control (RBAC), and security monitoring for local AI workflows.
User Enablement: Provide self-service environments (Jupyter, Kubeflow) for internal data science teams
Qualification & Role Requirements
-
Full time Engineering graduate.
- Core Stack: Deep Linux administration, Docker, and Kubernetes (K8s) orchestration.
- Hardware/Drivers: Hands-on experience with NVIDIA CUDA, NCCL, and GPU performance tuning.
- MLOps/LLMOps: Familiarity with model serving runtimes, MLflow, Kubeflow, or Red Hat OpenShift AI.
- Languages: Proficiency in Python and infrastructure automation tools (Terraform, Ansible).
- Background: 4 to 8+ years in DevOps, Site Reliability Engineering (SRE), or AI infrastructure operations