About the Role
We are looking for a DevOps Platform Engineer to design, size, and manage the cloud and infrastructure foundation for Generative AI and Agentic AI solutions, preferably within the Consumer Packaged Goods (CPG), Food & Beverage industry. This role owns the technical backbone that GenAI/Agentic applications run on — from infrastructure planning and BOM creation to GPU/compute sizing, LLMOps pipelines, and continuous optimization of AI workloads.
The ideal candidate understands both classic cloud/DevOps fundamentals and the unique infrastructure demands of LLMs, RAG pipelines, and multi-agent systems — and can independently translate a customer's Gen AI use case into a right-sized, production-ready, cost-optimized platform.
Key Responsibilities
Infrastructure & Requirements Assessment for Gen AI Workloads
- Engage with client stakeholders, data scientists, and solution architects to understand Gen AI/Agentic AI use cases (copilots, RAG applications, autonomous agents, document intelligence, etc.) and their infrastructure implications.
- Assess current-state client environments and identify gaps for hosting LLM-based and agentic applications (compute, GPU access, networking, data pipelines, security).
- Translate model requirements (model size, context length, throughput/latency targets, concurrency) into concrete infrastructure specifications.
BOM & Cloud Architecture for AI Platforms
- Prepare and maintain detailed Bill of Materials (BOM) covering compute (CPU/GPU), storage, networking, vector databases, orchestration tooling, and LLM API/licensing costs.
- Design cloud architecture (AWS/Azure/GCP) for Gen AI workloads — including model hosting/inference endpoints, vector databases, RAG pipelines, agent orchestration layers, and API gateways.
- Evaluate build-vs-buy decisions: managed LLM APIs (OpenAI, Anthropic, Azure OpenAI, Bedrock) vs. self-hosted/open-source models (Llama, Mistral, etc.) based on cost, data privacy, and performance needs.
- Support proposal and pre-sales efforts with accurate sizing, GPU costing, and architecture inputs for Gen AI engagements.
Sizing & Capacity Planning for LLM/Agentic Workloads
- Perform workload analysis and capacity planning specific to Gen AI systems — token throughput, concurrent users, embedding/indexing volumes, and agent execution loads.
- Size GPU/compute infrastructure for model inference and (where applicable) fine-tuning, balancing latency, throughput, and cost.
- Plan for vector database scale (embedding volume, query load) and retrieval pipeline performance for RAG-based solutions.
- Account for CPG-specific patterns — seasonal spikes (promotions, demand planning cycles), batch document/data processing volumes, and multi-brand/multi-market scaling needs.
Workload & Cost Optimization
- Monitor Gen AI infrastructure spend and performance — GPU utilization, API token consumption, inference latency, and vector DB query costs.
- Implement autoscaling and dynamic resource allocation for inference endpoints and agent workloads to manage cost-to-performance trade-offs.
- Apply FinOps practices tailored to AI workloads — cost attribution by use case/agent, model routing to lower-cost models where appropriate, caching, and prompt/token optimization strategies.
- Identify opportunities to right-size GPU instances, batch inference jobs, and reduce idle compute costs.
LLMOps / Platform Engineering & Operations
- Build and maintain infrastructure-as-code (Terraform, CloudFormation, ARM/Bicep) for repeatable provisioning of AI platform components.
- Set up CI/CD pipelines for Gen AI applications, including model/prompt versioning, evaluation gates, and safe rollout of agentic workflows.
- Deploy and manage container orchestration (Kubernetes/Docker) for model serving, RAG pipelines, and agent runtimes.
- Implement monitoring, logging, and observability for AI-specific metrics (latency, token usage, hallucination/error rates, agent task success rates) using tools like Prometheus, Grafana, Datadog, or LLM-specific observability platforms (e.g., LangSmith, Arize).
- Ensure infrastructure security, data privacy, and compliance for AI systems handling sensitive or proprietary client data.
Client & Stakeholder Engagement
- Act as the technical point of contact for infrastructure and cloud discussions on Gen AI/Agentic AI engagements.
- Present sizing, BOM, and architecture recommendations for AI platforms in clear, business-friendly terms to technical and non-technical stakeholders.
- Collaborate closely with Gen AI solution leads, data scientists, and application teams to ensure infrastructure choices align with use case goals and budget.
Required Qualifications
- 5–7 years of relevant experience in DevOps, Cloud Engineering, or Platform Engineering, with hands-on exposure to Gen AI/LLM infrastructure (inference hosting, RAG pipelines, agent frameworks, or MLOps/LLMOps).
- Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP), including AI/ML-specific services (SageMaker, Azure AI Foundry/OpenAI Service, Vertex AI, Bedrock).
- Experience sizing and costing GPU-based compute for model inference and/or fine-tuning workloads.
- Familiarity with vector databases (Pinecone, Weaviate, Milvus, pgvector, etc.) and RAG pipeline architecture.
- Working knowledge of LLM orchestration/agent frameworks (LangChain, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, or similar) — enough to understand infrastructure and integration implications, not necessarily to build agents.
- Practical experience with Infrastructure-as-Code (Terraform, CloudFormation, Ansible) and container orchestration (Docker, Kubernetes).
- Experience with CI/CD tooling (Jenkins, GitHub Actions, Azure DevOps, GitLab CI) adapted for AI/ML deployment pipelines.
- Understanding of FinOps principles applied to AI workloads — token cost management, GPU utilization tracking, model routing for cost efficiency.
- Experience working with or serving CPG / Food & Beverage industry clients strongly preferred.
- Strong communication skills with the ability to explain AI infrastructure trade-offs to client stakeholders and non-technical audiences.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field (or equivalent practical experience).
Preferred Qualifications
- Cloud AI/ML certifications (AWS Machine Learning Specialty, Azure AI Engineer Associate, GCP Professional ML Engineer).
- Experience deploying open-source/self-hosted LLMs (e.g., via vLLM, TGI, Ollama, or similar serving frameworks).
- Exposure to responsible AI/governance practices for infrastructure — access control, data residency, model auditing.
- Prior experience in a client-facing consulting or managed services environment delivering AI-enabled platforms.
- Familiarity with CPG-specific data ecosystems (SAP, Nielsen/IRI, plant/IoT systems) that commonly feed Gen AI use cases.
Pay: ₹6,483.30 - ₹20,952.10 per month
Work Location: Hybrid remote in Delhi, Delhi (Delhi)