AI Jobs Map

Altera Institute · Gurugram, Haryana, India

Artificial Intelligence Engineer

mid_levelfull timePosted yesterday
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

githubllmvllmpythonpytorchtransformersartificial-intelligencefine-tuninga/b-testingdeep-learninghugging-faceapi-designprompt-engineeringmachine-learning

AI Engineer - Inference & Optimization || Altera Institute - Digiaccel Learning

Note: Please share a link to a project, GitHub repository, technical blog, paper, demo, or other proof of work where you have worked hands-on with LLM fine-tuning, inference optimization, model evaluation, or GPU optimization. Briefly describe your contribution.

At Digiaccel Learning and Altera Institute, we are on a mission to redefine management education for the digital and AI-first world. We are building technology that powers transformative learning experiences for thousands of learners across India. From learner journeys and assessments to AI enabled learning experiences and internal platforms, our engineering team plays a critical role in shaping the future of education.

As an AI Engineer - LLM Inference & Optimization, you will work on building and optimizing real time Voice AI systems that power next-generation learning experiences. You will work directly with large language models, inference systems, evaluation frameworks and GPU infrastructure to improve model accuracy, latency, throughput and cost in production environments.

This is not an LLM wrapper or prompt-engineering role. This role is ideal for someone who understands how LLMs work under the hood, enjoys experimenting with models, measures performance rigorously and knows how to make models faster, smaller, and more reliable for production workloads.

Key Responsibilities

LLM Fine-Tuning & Model Engineering

• Fine-tune LLMs for production use cases using LoRA, QLoRA, PEFT, full fine tuning, distillation, or similar techniques.

• Fine-tune and distill smaller, faster models for narrow and repeated reasoning, classification, judgment, and other production tasks.

• Experiment with different model architectures, training approaches, datasets, and hyperparameters to improve model performance.

• Combine code and LLMs to build consistent, reliable, and production-ready AI systems.

• Analyse model behaviour and identify opportunities to improve accuracy, reliability, and consistency.

Inference Optimization & Performance

• Optimize LLM inference using quantization, batching, caching, streaming generation, and other performance optimization techniques.

• Work with GPU inference systems and optimize KV cache, GPU memory utilization, batching, and compute efficiency.

• Own model latency as a core engineering metric, including TTFT, p50/p95 latency, throughput, and tail latency.

• Optimize inference specifically for real-time Voice AI, where even a few hundred milliseconds can materially impact the quality of the conversation.

• Evaluate and optimize inference frameworks such as vLLM, SGLang, TensorRT-LLM, or similar technologies.

• Identify performance bottlenecks and implement solutions that improve system responsiveness, scalability, and cost efficiency.

Evaluation, Benchmarking & Experimentation

• Build and maintain robust LLM evaluation and benchmarking systems for production models.

• Curate evaluation datasets and design meaningful metrics to measure model quality and task performance.

• Run structured experiments across different model variants and fine-tuning approaches.

• Build regression testing frameworks to ensure model improvements do not introduce unexpected performance or quality degradation.

• Perform statistical comparisons between model variants and use data to guide model selection.

• Benchmark and evaluate trade-offs across accuracy, latency, model size, throughput, memory usage and cost.

Production AI Engineering

• Take ownership of AI models and inference systems from experimentation through production deployment and continuous optimization.

• Monitor production model performance and proactively identify opportunities for improvement.

• Collaborate closely with Product, Engineering, and other cross-functional teams to translate business problems into reliable AI solutions.

• Build systems that balance model intelligence with the practical requirements of speed, consistency, scalability, and cost.

• Contribute towards establishing an engineering discipline around model accuracy and latency across the AI platform.

What We're Looking For

• 3-5 years of experience in AI/ML engineering, deep learning, LLM engineering, or a closely related field..

• Strong programming skills in Python and hands-on experience with PyTorch and/or Hugging Face Transformers.

• Hands-on experience working with LLMs beyond API integration or prompt engineering.

• Practical experience with one or more of LoRA, QLoRA, PEFT, fine-tuning, or model distillation.

• Strong understanding of LLM inference, GPU computation, KV cache, batching, and memory optimization.

• Experience with quantization and inference optimization techniques.

• Exposure to vLLM, SGLang, TensorRT-LLM, or similar inference frameworks.

• Understanding of LLM evaluation, benchmarking, dataset curation, and model experimentation.

• Experience with low-latency or streaming inference systems is highly desirable.

• Strong understanding of the trade-offs between accuracy, latency, throughput, model size, memory, and cost.

• Strong analytical and problem-solving skills with an experimental and data-driven approach to engineering.

• Ability to work independently, take ownership of complex technical problems, and operate effectively in a fast-paced environment.

• Bachelor's or Master's degree in Computer Science, Engineering, Machine Learning, or a related technical discipline.

Why Join Us?

• Build Real-Time AI Systems - Work on production-grade Voice AI systems where model quality and latency directly impact the user experience.

• Work Beyond Prompt Engineering - Get hands-on exposure to model fine tuning, inference optimization, evaluation, GPU systems, and model performance.

• Solve Challenging AI Problems - Work on problems involving model accuracy, latency, scalability, memory, throughput, and cost.

• End-to-End Ownership - Contribute across the complete AI lifecycle, from experimentation and fine-tuning to evaluation, optimization, and production inference.

• High-Growth Environment - Be part of a fast-scaling organization that values innovation, speed, intellectual honesty, and engineering excellence.

• Continuous Learning - Work closely with experienced product and engineering leaders while building expertise in rapidly evolving AI technologies.

• Meaningful Impact - Build AI systems that power learning experiences and impact the careers and learning journeys of thousands of professionals.

About the Company

Digiaccel Learning is a Gurgaon headquartered education company that is redefining the gold standard of management education for the digital and AI first world. By doing so, it is building the next generation of business leaders in the country. The mission of the company is to build employability through education.

Education is a hard and important problem to solve for. India needs to double its higher education infrastructure in order to serve its large demographic of youth. Within what exists, quality is a large problem with studies reporting that 50% of Indian graduates are unemployable. Add to this the shift in the skills ecosystem at the workplace with AI and digitization. All of this has led to a higher-education infrastructure that is unable to serve the aspirations of its students. So, you have MBA students from legacy institutes who are unplaced and professionals who are struggling to level up in their careers because they don’t have the skills to progress.

With a learner-centric philosophy, industry-driven curriculum, experiential learning opportunities, and deep engagement with corporate leaders, Altera equips learners with the skills, capabilities, and mindset required to thrive in today's rapidly evolving business landscape.

We are looking for smart passionate folks to join us in this mission of building a better, smarter and more outcome-oriented education system. We are a young company and each role has the ability to make a large impact. We can assure of a very talent dense, intellectually honest, action oriented and collaborative environment.

We are sure that you will join for the opportunity and stay for the impact of seeing your work help thousands of others.

More jobs at Altera Institute