Role: AI Lead
Location: Noida
Type: Full-Time, Permanent
Experience: 8+ years
Role Overview
We are looking for a hands-on Technical Program Lead to design and deliver enterprise
GenAI/SLM solutions, including air-gapped, on-prem, and sovereign deployments. You will
own architecture end-to-end — model selection, infra, deployment, and governance — while
leading delivery and the client relationship.
Key Responsibilities
● Architect GenAI/SLM solutions (RAG, agentic workflows, fine-tuning/distillation) suited to
customer security and data-sensitivity constraints.
● Evaluate SLMs vs. LLMs (Phi, Mistral, Llama, Qwen, etc.) on cost, latency, and accuracy
trade-offs.
● Design air-gapped/offline deployments — local inference, vector stores, and secure
model/data update pipelines with no external dependency.
● Architect across hybrid environments: AWS/Azure/GCP, private cloud, and on-prem data
centers, optimizing GPU/CPU cost and performance.
● Define AI governance: model evaluation, guardrails, audit logging, and responsible-AI practices — including offline-compatible monitoring for restricted environments.
● Lead client discovery workshops, translate business requirements into a scoped delivery
roadmap, and drive the engagement through to shipment/go-live.
● Own planning and task allocation across the team — break architecture into workstreams, assigned to the right engineers, and sequence delivery against client timelines.
● Be the primary point of client interaction throughout the engagement — status updates, scope changes, escalations — not just at kickoff/handoff.
● Drive multiple projects/accounts in parallel, balancing priorities across engagements and flagging capacity or scope risk early.
● Lead a team of engineers/data scientists — planning, reviews, and unblocking delivery.
● Support pre-sales: scoping, estimation, and technical proposals.
Required Skills & Experience:
● 8+ years in software/data engineering, 3+ years architecting production ML/GenAI solutions.
● Hands-on with SLMs/LLMs, fine-tuning (LoRA/QLoRA), quantization; Python, LangChain/LlamaIndex, vLLM/Ollama.
● Proven experience with air-gapped or on-premise AI deployment.
● Cloud architecture (AWS/Azure/GCP) plus hybrid/private data center deployment.
● Vector DBs deployable offline (FAISS, Milvus, Weaviate, Qdrant).
● Familiarity with AI governance/compliance frameworks (NIST AI RMF, ISO/IEC 42001) and data residency requirements.
● Docker/Kubernetes and infra-as-code (Terraform/Ansible).
● Expert in Claude-driven development — using Claude Code and Claude-based agents as a core part of the build workflow, including authoring custom Skills/MCP tools and agentic coding pipelines to boost team engineering productivity.
● Reviewer-first mindset: with agents doing most of the generation, your value is in specifying correctly, critically reviewing AI-generated architecture/code, catching subtle design and security flaws, and validating trade-offs — not in hand-writing every line yourself.
Behavioural & Leadership Expectations:
● Must have: prior experience leading a small team (formally or as a de facto lead) and
working across multiple clients/engagements simultaneously — this is not a first team-lead or first multi-client role.
● Leads a team end-to-end; owns the client relationship from requirement gathering through shipment.
● Spends more time planning, allocating, and reviewing than hand-coding — sets direction, defines specs/guardrails for agentic tooling, allocates tasks across the team, and audits output; comfortable being judged on decision quality and delivery outcomes, not lines of code written.
● Able to run multiple projects/accounts simultaneously without losing quality of client interaction on any one of them.
● Fluent in Agile/Scrum ceremonies; hands-on with JIRA/Confluence for backlog and delivery tracking.
● Self-driven, strong client-facing communicator across technical and non-technical stakeholders.
● Preferred: background in an IT/consulting services company over purely captive/product environments.
Good to Have
● Big Data (Spark/Hive/Hadoop), Graph Analytics, or hardware acceleration (GPU/FPGA) experience.
● Regulated-industry (defense, government, BFSI) AI deployment experience.
● Cloud, security, or AI governance certifications.
● Contribution to open source projects, academic papers published, filled patents.