AI Jobs Map

Flam · Bengaluru, Karnataka, India

Artificial Intelligence Engineer

mid_levelfull timePosted today
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmvllmobservabilitypythonpytorchnlpgcpawsonnxmilvusredisdockerartificial-intelligencefine-tuning

Flam is building the next generation of interactive media through its content format. We are an AI-native technology company transforming how brands and consumers interact through immersive, interactive content. Our technology enables rich, app-less experiences that can be launched instantly on smartphones, creating a fundamentally different way for brands to engage consumers. We are backed by leading investors and already work with some of the world's largest brands. We are now building Flicks, our interactive media format for the US market.

About the role

Flam builds multimodal AI systems that ship to real customers. Our stack spans four production surfaces:

Falcon — Our LLM system, built on a sparse-MoE backbone.

Finesse — Our TTS and cross-lingual voice cloning system.

We are a small team. Whoever joins will own systems end-to-end training, evaluation, serving, and the latency budget.

Responsibilities:

Fine-tune and evaluate open-weight models LoRA/QLoRA, full SFT, preference tuning) and get them into production without a quality regression.

Optimize inference: quantization FP8/NVFP4/INT4, speculative decoding, prefix and KV-cache strategies, batching and scheduling on vLLM or SGLang.

Build evaluation harnesses that tell us something true — benchmark suites, regression gates, and per-release comparisons against both our own prior checkpoints and external baselines.

Own latency. Profile the pipeline, find where the milliseconds go, and remove them.

Take research to production: read the paper, replicate it, decide honestly whether it's worth shipping, and then ship it.

Write and maintain the serving infrastructure around your models —containers, autoscaling, GPU scheduling, observability.

What we're looking for

Required

2+ years building ML systems that ran in production, not only in notebooks.

Strong Python and PyTorch. You can read a model implementation and modify it, not just call .fit.

Hands-on experience with at least one modern inference stack (vLLM, SGLang, TensorRTLLM, or TGI) and a real understanding of what makes it fast.

Demonstrated fine-tuning experience — you've trained something, evaluated it properly, and know why your eval numbers meant what you claimed.

Comfort with GPU-level reasoning: memory layout, precision trade-offs, where the bottleneck actually is.

Ability to work from a paper. We move on recent research and expect you to be able to read it.

Strongly preferred

Experience in one or more of: speech ASR/TTS, diffusion and image generation, or video/avatar generation.

Quantization experience beyond running an off-the-shelf script.

Real-time streaming systems WebRTC, low-latency audio/video pipelines).

Indic language NLP, multilingual tokenizer work, or code-mixed data.

Cloud GPU deployment on GCP, AWS, or RunPod — including the unglamorous parts.

Our stack

PyTorch · vLLM · SGLang · TensorRT / ONNX · Unsloth · LLaMAFactory · TRL Milvus · BGEM3 · Redis · NATS · Docker · GCP · RunPod · WebRTC · ComfyUI.

More jobs at Flam