About The Role
The LLM / GenAI Engineer will design, build, and operate production AI systems built on foundation models, including retrieval-augmented generation, tool-using agents, model adaptation, and automated evaluation. The role spans experimentation and backend engineering, with responsibility for turning promising prototypes into reliable services with measurable quality, latency, and cost targets.
The engineer will work with applied scientists, product engineers, and platform teams to integrate LLM capabilities into customer-facing workflows. This role is based in Phoenix, AZ and operates remotely, with a focus on secure data handling, observability, reproducible deployments, and continuous improvement of model behavior in production.
Key Responsibilities
- Design and implement RAG systems using Python, LangChain, LlamaIndex, or custom orchestration services, including ingestion, chunking, retrieval, reranking, and citation workflows
- Build production integrations with vector stores such as Pinecone, Weaviate, OpenSearch, or pgvector, optimizing embedding quality, retrieval accuracy, throughput, and latency
- Develop agentic workflows that safely call internal APIs and business tools, with structured outputs, permission controls, retries, tracing, and human-in-the-loop escalation
- Create LLM evaluation frameworks covering offline benchmarks, task-specific quality metrics, hallucination and safety checks, LLM-as-judge pipelines, and regression testing
- Adapt foundation models through prompt optimization, supervised fine-tuning, and parameter-efficient methods such as LoRA or QLoRA using PyTorch, Hugging Face, or equivalent tooling
- Deploy and operate inference services on AWS, Azure, or GCP using Docker, Kubernetes, and CI/CD pipelines; monitor token usage, latency, errors, drift, and service reliability
- Collaborate on architecture reviews and write clean, tested, documented code that meets security, privacy, and production-readiness requirements
What We Are Looking For
- 3–8 years of software engineering, machine learning engineering, or applied AI experience, including at least 1 year delivering LLM-powered systems to production
- Strong Python skills with experience building asynchronous services, REST or gRPC APIs, data pipelines, and automated tests
- Hands-on experience with LLM application patterns including RAG, embeddings, vector search, prompt engineering, structured generation, and agent orchestration
- Proficiency with at least one major LLM or ML ecosystem such as OpenAI or Anthropic APIs, Hugging Face Transformers, PyTorch, vLLM, LangChain, or LlamaIndex
- Experience deploying cloud-based services with Docker and Kubernetes, plus practical knowledge of observability tools, CI/CD, secrets management, and API security
- Bachelor’s or master’s degree in computer science, machine learning, artificial intelligence, data science, or a related technical field, or equivalent professional experience
- Bonus: Experience with fine-tuning open-weight models, multimodal models, distributed inference, MLflow, Ray, Terraform, guardrail frameworks, or privacy-sensitive enterprise data