Principal AI Engineer
Location: India (Bengaluru) - 3(WFO)
Employment Type: Full Time (Overlapping EST)
Experience Level: Staff/Principal (8–14 years)
What We're Looking For
Engineering foundation
- 8–14 years of software engineering experience, with strong hands-on large-scale Python
- Working depth in at least one systems or backend language — Go, Rust, Java, or C/C++ — and the judgment to know when to reach for it
- Strong data structures and algorithms.
- Strong understanding of APIs, microservices, and system design
- Hands-on experience building and operating data pipelines and production-grade distributed systems.
Agentic AI and LLMs
- 2+ years of hands-on LLM engineering, with at least couple agentic system you designed and took to production
- Production experience with agent frameworks — LangGraph, Google ADK, CrewAI, Claude Agent SDK, or equivalent — and the fluency to move between them as the ecosystem evolves
- Experience building MCP (Model Context Protocol) servers and tool-calling interfaces
- RAG from first principles: chunking strategy, embeddings, vector and hybrid retrieval, reranking, and response validation
- Strong experience with vector databases (Milvus, Pinecone, Weaviate, FAISS, etc. or cloud equivalents)
- Design of guardrails and reliability patterns — validators, policy checks, self-correction loops, deterministic fallbacks, circuit breakers, and rollback paths
-
Optimization
- Deep familiarity with token optimization and context-window management — context shaping, pruning, and compaction
- Latency and cost optimization through caching, model routing, batching, streaming, and parallel tool calls
- Performance testing and tuning systems against defined SLOs
Evaluation
- Experience building evaluation frameworks for LLM systems — offline eval sets, continuous online evaluation, and regression detection
- Instrumentation and traceability suitable for regulated enterprise environments using tools like LangSmith, Langfuse, etc.
Cloud
- Hands-on AWS: containerized services (ECS/EKS), serverless (Lambda), data services (S3, DynamoDB, Redshift) and orchestration (Step Functions)equivalents also valued
- Familiarity with CI/CD pipelines and DevOps practices
- Infrastructure as code with Terraform or CloudFormation, and mature CI/CD practice
Working traits
- Strong analytical problem-solving with a bias to ownership and urgency
- Clear cross-team communication, working directly with client stakeholders to translate business problems into technical roadmaps
- Able to work productively in ambiguity from system-level documentation and ramp quickly in unfamiliar codebases
Good to Have
- Experience with managed AI platforms — Amazon Bedrock, Vertex AI, Azure AI — paired with fluency in the underlying fundamentals
- Azure or GCP
Roles & Responsibilities
- Design and build agentic systems: Lead the architecture and implementation of tool-calling agents that combine retrieval, structured reasoning, and secure action execution with least-privilege access.
- Productionize LLM applications: Build retrieval pipelines, prompt synthesis, response validation, and self-correction loops, backed by rigorous evaluation.
- Own the full stack: Deliver the data pipelines, backend services, distributed compute, and orchestration layer that agentic systems depend on — not only the model invocation.
- Engineer for reliability and governance: Build validator models, adversarial test suites, and policy checks; enforce deterministic fallbacks and rollback strategies; instrument continuous evaluation.
- Optimize for cost and latency: Drive measurable improvements in token efficiency, response time, and unit economics against defined SLOs.
- Codebase ownership: Build, maintain, and review high-quality Python and SQL, with an emphasis on reusable components, scalability, and performance.
- Cloud integration: Deploy AI applications on AWS, Azure, or GCP with optimized resource usage and robust CI/CD.
- Cross-functional collaboration: Partner with product owners, data scientists, and business SMEs to define requirements and deliver impactful AI products.
- Mentoring and technical leadership: Set engineering standards and share knowledge across the team, raising the bar on AI and software engineering practice.