AI Jobs Map

BCI~IT · Delhi, India

Sr. GenAI / Python / LLM Fine Tuning Engineer

Remoteseniorfull timePosted 16 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

generative-aipythonllmfine-tuningreinforcement-learninggeminiprompt-engineeringapi-designawsgcppytorchhugging-facetransformersvector-databasesragmlopsmlflowkubeflowvllmnlp

BCI has an open position on our offshore GenAI team working with our USA based client. The GenAI / Python / LLM Fine Tuning Engineer will join our offshore development team that is growing and there is a lot of new and exciting GenAI work to be completed. This is a full-time remote position and must be able to work blended hours of EST / IST timings. *Please note: In addition to doing GenAI / Python Development, this role will also focus on LLM Fine tuning - LoRA / QLoRA, PEFT, full fine-tuning, instruction tuning, RLHF/DPO. Fine-tune and adapt large language models (Llama, Gemma, and other open-weight models.*

About the Role

We're looking for an GenAI / Python / LLM Fine Tuning Engineer for the customization and optimization of large language models for production use cases. This role involves full fine-tuning lifecycle — from data preparation through training, evaluation, and deployment — working with open-weight models (e.g., Llama, Gemma) as well as proprietary/managed models (e.g., Google Gemini) where fine-tuning access is available. You'll be hands-on with real training runs at scale, not just prompt engineering, or API integration.

What You'll Do

- Fine-tune and adapt large language models (Llama, Gemma, and other open-weight models, plus managed options like Gemini where applicable) for specific business use cases

- Design and execute full fine-tuning pipelines: dataset curation and cleaning, tokenization, training/eval splits, hyperparameter selection, and training runs (full fine-tune, LoRA/QLoRA, PEFT, RLHF/DPO as appropriate)

- Run and manage large-scale training jobs across multi-GPU / distributed environments

- Evaluate model performance using both automated benchmarks and human-in-the-loop review; iterate to close quality gaps

- Optimize models for production inference (quantization, distillation, latency/cost tradeoffs)

- Deploy fine-tuned models into production systems and monitor performance, drift, and degradation over time

- Collaborate with data, ML infrastructure, and product teams to define fine-tuning objectives and success metrics

- Stay current on the open-model landscape and evaluate new base models as candidates for fine-tuning

- Document methodology, training runs, and results for reproducibility and knowledge sharing

Required Qualifications

- Strong proficiency in Python, with solid software engineering fundamentals (not just notebooks)

- Hands-on, production-level experience GenAI and fine-tuning LLMs — this is a must-have, not exploratory/academic experience only

- Demonstrated experience taking fine-tuned models into live, large-scale production systems (not just POCs)

- AWS and /or Google cloud production experience will be considered

- Experience with open-weight model families (e.g., Llama, Gemma, Mistral, or similar)

- Practical knowledge of fine-tuning techniques: LoRA/QLoRA, PEFT, full fine-tuning, instruction tuning, RLHF/DPO

- Experience with ML/training frameworks such as PyTorch, Hugging Face Transformers/TRL/PEFT, DeepSpeed, or similar

- Familiarity with distributed/multi-GPU training and the associated infrastructure challenges

- Solid understanding of model evaluation methodology for generative models

Nice to Have

- Experience fine-tuning or customizing Google Gemini or other managed/API-based models

- Experience with vector databases, RAG architectures, or hybrid RAG + fine-tuning approaches

- Experience with MLOps tooling for training pipelines (e.g., MLflow, Weights & Biases, Kubeflow, SageMaker, Vertex AI)

- Experience with model quantization and inference optimization (vLLM, TensorRT-LLM, GGUF, etc.)

- Background in NLP research or publications related to LLM training/fine-tuning

What Success Looks Like

Within your first few months, you're independently running fine-tuning jobs on open models, have a clear point of view on which base models and techniques fit which use cases, and have shipped at least one fine-tuned model into a production system with measurable quality improvement over baseline.

Interview Process:

1. If profile appears to fit role, we will send you a request for more information and details on your background. 2. Initial 30 min MS Teams conversation with BCI-IT team to go over your hands on experience and determine fit. 3. If potential fit, you will be sent a video technical screen with several questions on LLM fine tuning. 4. 45 min to 1 hour client technical interview with code share activity and technical discussion. You will be speaking with 2-3 Sr. team members. Hiring decision can be made after call.

More jobs at BCI~IT