Company: AutoVRse
Location: Bangalore
Work Mode: On-site/work from office
Experience: 1-3 years
Employment Type: Full-time
Salary: Up to 12 lacs
About AutoVRse
AutoVRse is a Bengaluru-based industrial XR and AI company building VR Training, AR/AI solutions and Smart Glasses based solutions for enterprise clients in manufacturing, pharmaceuticals, oil & gas, and allied sectors. We work with some of the largest industrial organizations in India and globally, and we're scaling rapidly off a recent funding round. Since 2016 we have run 250+ deployments, worked with 100+ enterprise clients, and trained over 300,000 workers across manufacturing, pharma, energy, oil and gas.
Learn more about us here: About AutoVRse
About the Role
We are building AI products for real industrial workflows - systems that have to make sense of video, images, audio, documents and structured data together, and produce reliable, grounded output that a plant team can actually act on.
We already have strong engineers who know how to build and scale production software: a Unity-based product used by external developers, and Node.js and React services running at scale on AWS. What we are adding is a different kind of engineer - someone who has spent the last few years deep inside applied AI, has built multiple AI systems, watched some of them fail, worked out why, tried alternative architectures, compared models, and developed an instinct for which approach is appropriate for which problem.
This is not a role for someone whose AI experience is limited to calling LLM APIs. A typical problem here is not “which prompt should we send to GPT?”. It is closer to: “given these SOP documents, plant videos, inspection photos and field recordings, what combination of parsing, retrieval, embeddings, vision models, LLMs and deterministic processing gives us the most reliable result?” Sometimes the correct answer will be an LLM. Sometimes it will be a VLM, an embedding model, a reranker, a classical ML or vision technique, or an entirely different architecture. We expect you to know the difference - and to be able to say “not with this data, not yet” when that is the honest answer. You will own AI work end to end, including its integration into our existing services, and you will have real influence on what we build.
Key Responsibilities
- Build and own AI systems end to end - problem framing, retrieval and model design, evaluation, deployment, and integration into our Node.js and React services.
- Design, test and rebuild retrieval architectures - chunking and indexing strategies, dense vs sparse vs hybrid retrieval, reranking, multimodal and cross-modal retrieval - and understand why each one breaks.
- Work on multimodal understanding across video, images, audio and documents: knowledge extraction, structured information generation, grounding and citations, and hallucination reduction.
- Select and benchmark models rather than defaulting to them - LLMs, VLMs, embedding models, rerankers, vision and OCR models, speech models and smaller specialist models - and make the case for the boring option when it is the right one.
- Build evaluation before features: ground-truth sets, benchmarks, failure taxonomies and regression checks.
- Optimise inference for latency and cost, and design agentic / multi-stage workflows and model routing where they genuinely earn their complexity.
- Look at a new customer data source and quickly identify what is technically possible, what is not, and what would need to be in place first - with your confidence level stated.
- Prototype fast, discard fast, and document what you learned.
- Raise the AI engineering capability of the wider team. You will be explaining and teaching as much as building.
Required Skills
- 1-3 years of hands-on applied AI experience, with multiple AI projects that reached real users.
- Excellent Python and fluency with the surrounding ecosystem: PyTorch, Transformers / Hugging Face, NumPy, Pandas, vector databases, LlamaIndex / LangChain or equivalent, and model APIs and inference SDKs.
- Multiple retrieval approaches built and compared - dense vs sparse vs hybrid, embedding model selection for a given dataset, when reranking materially helps, how to index large documents or long videos - with an opinion on where each one fell over, and a view on when RAG is the wrong architecture entirely.
- A clear, practical understanding of what LLMs are good at, what their output quality actually depends on, and where a non-generative model is simply the better tool.
- Breadth beyond the heavily marketed AI stack - open-source models, vision and multimodal embedding models, object detection and image understanding, speech and audio models, OCR and document intelligence, and newly published architectures and retrieval techniques.
- Production experience - AI services and APIs you have shipped and that were actually used, deployment on cloud infrastructure, async and distributed processing, queues and background jobs, observability, and handling large, messy real-world datasets.
- Comfort going from a paper, model card or repository to a working prototype within days.
Good to Have
- A multimodal project involving video, images or audio.
- Fine-tuning, adapting, quantising or locally running open-source models - especially getting a smaller model to beat a larger one on a narrow task.
- Computer vision experience, or work involving multimodal embedding spaces.
- Open-source contributions, technical blogs, papers or detailed engineering write-ups.
- Exposure to AR/VR, Unity or 3D applications.
- Any experience with industrial, manufacturing or field-operations environments.
What We're Looking For
We value curiosity, ownership, experimentation and evidence over credentials. You are probably a strong fit if:
- You experiment with new AI models and papers because you genuinely enjoy it, not because someone asked you to.
- Your GitHub, notebooks or side projects contain things you built purely because you wanted to understand how they worked.
- You have strong opinions about when certain systems perform well vs bad.
- You enjoy ambiguous problems where the solution architecture is not already known.
- You move fast and you take ownership rather than waiting for a perfectly specified task.
Work Environment
This is an on-site role based out of our Bangalore office. You will work closely with multidisciplinary teams across AI, software engineering, 3D, AR/VR and product development.
Why Join AutoVRse?
- You will bring the AI depth on the team - your model and architecture decisions will materially shape the product rather than implement someone else's roadmap.
- Problems from real enterprise deployments, not isolated research exercises: incomplete data, long videos, noisy field recordings, complex documents, domain-specific knowledge and strict accuracy requirements.
- Genuine freedom to investigate better approaches, challenge existing architectures and introduce technologies the rest of the team may not yet know about.
- Work at the intersection of AI and immersive technology, in a company that already has enterprise distribution and production software running at scale.
- A small, fast-moving team and a very short distance between your work and a customer using it.
How to Apply
If you or anyone you know feels like joining us for this ride - please feel free to reach out to us at: [email protected] /[email protected]