Job Title: Data Scientist
Work Location: Bangalore / Bengaluru (Hybrid)
Position Type: Permanent
Interview: Virtual + F2F
Required Experience: 5 years
Need Immediate joiners max 30 days notice
Max Salary: 24LPA
Must have skills: Data Scientist-5+ years, RAG architecture, API integrations, LLM Ops tools, R&D applications
Job Description:
As a Data Scientist, you will identify business trends and solve complex problems
using large-scale data and advanced AI techniques. You will design, develop, and
deploy high-impact solutions ranging from classical ML/DL to LLM-powered
applications including RAG-based architectures and Agentic AI systems (tool-using,
workflow-driven AI that can plan, reason, and act). You will collaborate with
stakeholders and cross-functional teams to drive innovation, improve existing
intelligent products, and deliver measurable outcomes.
Key Responsibilities
- Product & Model Optimization: Analyze existing digital products to improve the performance, reliability, and scalability of current intelligent models.
- LLM Integration: Enhance traditional machine learning (ML) and deep learning (DL) pipelines with LLM features like summarization, Q&A, reasoning, decision support, and copilots.
- RAG Architecture: Design and deploy Retrieval-Augmented Generation (RAG) solutions grounded on enterprise data (documents, manuals, telemetry, tickets, knowledge bases).
- Agentic Workflows: Build autonomous AI workflows with multi-step task planning, function calling for tools/APIs, guardrails, approvals, and contextual memory management.
- Orchestration & Fallbacks: Develop multi-agent collaboration patterns (planner-executor-critic), deterministic workflow engines, and fallback strategies for low-confidence outputs.
- R&D & Patents: Drive innovation through rapid experimentation, patent filings, invention disclosures, and novel solution design.
- Specialized Domains: Implement targeted AI solutions for IoT, robotics, and industrial automation use cases.
- Pipeline Engineering: Construct scalable pipelines for training, evaluation, and deployment across batch and real-time inference environments.
- LLMOps & Governance: Oversee experiment tracking, model registries, and versioning to guarantee reproducibility, traceability, and compliance.
- LLM Evaluation: Define and monitor generative metrics including groundedness, faithfulness, hallucination rate, toxicity, and safety.
- Retrieval Quality: Track vector search performance using metrics like precision, recall, chunking effectiveness, latency, and knowledge coverage.
Required Qualifications
- Education: Master’s degree in Computer Science, Electrical Engineering, Applied Mathematics, Statistics, or a related field (PhD preferred).
- Communication: Exceptional oral and written skills; capable of explaining complex technical concepts to non-technical stakeholders.
- Problem Solving: Proven ability to translate ambiguous business objectives into innovative, flexible solutions.
- Execution: Demonstrated track record of driving change and delivering impactful outcomes in complex environments.
Required Technical Skills
- Core LLMs: Deep expertise in developing scalable, production-grade LLM applications.
- RAG Systems: Hands-on experience designing and deploying end-to-end RAG architectures.
- Data Processing: Experience building ingestion and preprocessing pipelines for unstructured and semi-structured data.
- Chunking Strategies: Expertise in defining chunking methods to maximize retrieval precision and context relevance.
- Vector Mechanics: Strong understanding of embeddings, vector representations, and semantic search algorithms.
- Retrieval Optimization: Experience implementing hybrid search, dense retrieval, and reranking mechanisms to increase response accuracy.
- Explainability: Knowledge of grounding techniques and inline citation strategies for transparent LLM responses.
- Evaluation Frameworks: Hands-on experience establishing test harnesses to measure RAG quality and runtime performance.
- Tool-Using Agents: Experience building autonomous agents utilizing function calling and external API integrations.
- Agent Workflow Design: Proven ability to implement controlled, safe execution patterns for multi-step agent actions.
Preferred / Value-Add Skills
- Vector Databases: Hands-on knowledge of tools like Pinecone, Milvus, Weaviate, Elasticsearch/OpenSearch, Azure AI Search, or FAISS.
- Agent Frameworks: Familiarity with LangChain, LlamaIndex, Semantic Kernel, or deterministic workflow orchestration engines.
- LLMOps Tooling: Exposure to prompt management, version control, observability platforms, A/B testing, and red-teaming.
- Industrial R&D: 3+ years of experience in R&D, supported by published papers, patents, or patent applications.
- Domain Expertise: 3+ years of applied experience in:
- Robotics and automation (including Reinforcement Learning)
- Optimization theory (including black-box optimization)
- Designing IoT algorithms tailored for resource- and power-constrained hardware
- Cloud & DevOps: Experience with AWS, Azure, or GCP, along with containerization (Docker) and Kubernetes orchestration.
Behavioral Competencies
- Ownership: Proactive, self-driven mindset focused on discovering opportunities and taking end-to-end initiative.
- Pragmatism: Strong capacity to balance competing priorities and ship practical solutions efficiently.
- Collaboration: Highly collaborative approach paired with an innovation-first mindset.