About The Role
The LLM / GenAI Engineer will design, build, and deploy production AI systems, including retrieval-augmented generation pipelines, tool-using agents, model fine-tuning workflows, and evaluation infrastructure. The role spans experimentation and backend engineering, with a focus on reliability, latency, cost, and measurable model quality.
Working with applied scientists, platform engineers, and product teams, the role will turn emerging foundation-model capabilities into secure, observable services used in real customer workflows. The position is remote, with a preference for candidates based in Austin, TX.
Key Responsibilities
- Design and implement production RAG and agentic workflows using Python, LangChain, LlamaIndex, or equivalent frameworks
- Build ingestion, chunking, embedding, reranking, and retrieval pipelines backed by vector stores such as pgvector, Pinecone, Weaviate, or Milvus
- Develop LLM evaluation systems covering offline benchmarks, groundedness, factuality, safety, latency, cost, and regression testing
- Fine-tune and optimize open-source language models using supervised fine-tuning, LoRA, QLoRA, quantization, and distributed training techniques
- Deploy model-powered services through containerized APIs and cloud infrastructure using Docker, Kubernetes, AWS, GCP, or Azure
- Instrument production systems with tracing, logging, feedback collection, and monitoring for quality degradation, drift, latency, and token usage
- Partner with software engineers and applied scientists on architecture reviews, data strategy, experimentation, code quality, and production incident response
What We Are Looking For
- 3–8 years of experience in software engineering, machine learning engineering, or applied AI, including at least 1 year delivering LLM or GenAI systems to production
- Advanced Python skills with experience building asynchronous services, REST or gRPC APIs, testing frameworks, and maintainable production code
- Strong understanding of transformer-based language models, embeddings, tokenization, context windows, prompting, structured generation, and inference tradeoffs
- Hands-on experience with RAG architecture, vector databases, hybrid search, reranking, document processing, and retrieval-quality measurement
- Experience with at least one major cloud platform and production deployment tooling, including Docker, Kubernetes, CI/CD, and observability systems
- Bachelor’s or master’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent practical experience
- Bonus: Experience with distributed GPU training or inference, open-source model serving, multimodal models, guardrails, MLflow, vLLM, TensorRT-LLM, or enterprise AI security