About the Role
We're hiring an AI Engineer to deploy and productionize large open-weight LLMs entirely on-premise, in an air-gapped environment with no external API dependencies. You'll own the full stack from GPU infrastructure and inference serving up through the LLM-backed features that end users actually touch — assisted review, natural-language querying over structured data, and bulk triage workflows. This role is for someone who wants to solve real inference-engineering problems (throughput, latency, memory) rather than wrap a hosted API.
What You'll Do
- Deploy and operate large open-weight LLMs on-premise across multi-GPU, H100-class node infrastructure.
- Stand up and tune inference serving stacks (vLLM, TensorRT-LLM, or equivalent) for production workloads.
- Apply quantization, batching, and capacity sizing strategies to support concurrent users within hardware constraints.
- Design and build LLM-backed application features on top of the serving layer, including:
- Assisted review workflows
- Natural-language querying over structured data
- Bulk triage tooling
- Operate entirely within an air-gapped / restricted-network environment, with zero reliance on external APIs or cloud LLM providers.
- Monitor and optimize inference performance, GPU utilization, and cost/throughput trade-offs over time.
- Collaborate with infrastructure and product stakeholders to translate feature requirements into deployable, resource-aware LLM systems.
Must-Have Qualifications
- Demonstrated experience deploying large open-weight LLMs on GPU infrastructure on-premise (multi-GPU, H100-class nodes).
- Hands-on experience with inference serving and optimization — vLLM, TensorRT-LLM, or equivalent — including quantization, batching, and sizing for concurrent users.
- Experience building LLM-backed features, such as assisted review, natural-language querying over structured data, or bulk triage workflows.
- Comfortable working in an air-gapped / restricted-network environment with no external API dependencies.
Nice to Have
- Experience with model quantization formats (GPTQ, AWQ, GGUF) and trade-offs between them.
- Background in evaluating and benchmarking open-weight models for domain-specific tasks.
- Prior work in regulated, government, or otherwise network-restricted environments.
Pay: ₹477,682.82 - ₹1,780,277.15 per year
Work Location: Hybrid remote in Ahmedabad, Gujarat 380054