Data Scientist
Experience:5 years+
Location: Bangalore (Hybrid)
As a Data Scientist, you will identify business trends and solve complex problems using large-scale data and advanced AI techniques. You will design, develop, and deploy high-impact solutions ranging from classical ML/DL to LLM-powered applications including RAG-based architectures and Agentic AI systems (tool-using, workflow-driven AI that can plan, reason, and act). You will collaborate with stakeholders and cross-functional teams to drive innovation, improve existing intelligent products, and deliver measurable outcomes.
Key Responsibilities:
- Analyze existing digital products to understand current intelligent models and improve their performance, reliability, and scalability.
- Enhance traditional ML and DL pipelines by incorporating LLM-based capabilities such as summarization, Q&A, reasoning, decision support, and copilots.
- Design and implement LLM-based solutions using Retrieval‑Augmented Generation (RAG) to ground responses on enterprise data including documents, manuals, telemetry, tickets, and knowledge bases.
- Build Agentic AI workflows that enable multi-step task planning, tool and API invocation through function calling, controlled action execution with guardrails and approvals, and contextual memory management.
- Develop agent orchestration patterns such as multi-agent collaboration (planner–executor–critic), deterministic workflow engines, and fallback mechanisms for low-confidence retrieval or reasoning.
- Drive innovation through experimentation and contribute to invention disclosures, patents, and novel solution approaches.
- Design and implement AI solutions for IoT, robotics, and automation use cases.
- Build and maintain scalable pipelines for model training, evaluation, and deployment across batch and real-time inference scenarios.
- Manage experiment tracking, model versioning, and model registries to ensure reproducibility, traceability, and governance.
- Define and track LLM-specific evaluation metrics, including groundedness, faithfulness, hallucination rate, toxicity, and safety.
- Monitor retrieval system quality using metrics such as precision, recall, chunking effectiveness, latency, and knowledge coverage.
Required Qualifications:
- Master’s degree in Computer Science, Electrical Engineering, Applied Mathematics, Statistics, or a related field (PhD preferred).
- Strong oral and written communication skills; ability to explain technical concepts to non-technical stakeholders.
- Demonstrated ability to take ambiguous objectives and design innovative, flexible solutions.
- Proven track record of delivering impactful outcomes and driving change in complex environments.
Required Technical Skills:
- Strong expertise in Large Language Models (LLMs) and building scalable, production-grade applications using them.
- Hands-on experience designing and implementing Retrieval‑Augmented Generation (RAG) architectures.
- Experience building document ingestion and preprocessing pipelines for unstructured and semi-structured data.
- Expertise in defining effective chunking strategies to optimize retrieval quality and context relevance.
- Strong understanding of embeddings, vector representations, and vector search techniques.
- Experience implementing retrieval and reranking mechanisms to improve response accuracy.
- Familiarity with grounding and citation strategies to ensure reliable and explainable LLM outputs.
- Hands-on experience establishing evaluation frameworks to measure RAG quality and performance.
- Experience building tool-using agents leveraging function calling and API integrations.
- Proven ability to design and implement multi-step agent workflows with safe and controlled execution patterns.
Preferred / Nice-to-Have Skills (Strong Value Add):
- Experience with vector databases and search platforms (e.g., Pinecone, Milvus, Weaviate, Elasticsearch/OpenSearch vector, Azure AI Search, FAISS).
- Familiarity with agent frameworks/orchestration (e.g., LangChain, Semantic Kernel, LlamaIndex) and workflow engines for controlled execution.
- Experience with LLMOps tooling: prompt/version management, evaluation harnesses, observability, A/B testing, red teaming.
- 3+ years of industrial R&D with publications/patents/patent applications.
- 3+ years experience in:
- robotics/automation (including reinforcement learning),
- optimization theory (including black-box optimization),
- designing IoT algorithms under resource/power constraints.
- Cloud experience (Azure/AWS/GCP), containerization (Docker), and scalable deployment patterns (Kubernetes).
Behavioral Competencies:
- Strong ownership mindset; proactive in identifying new opportunities and leading initiatives.
- Ability to reconcile competing priorities and deliver pragmatic solutions.
- Collaborative team player with an innovation-first approach.