About the role
We’re looking for an AI Engineer to build production AI capabilities using both commercial AI APIs and open-source models.
You will develop LLM applications, build retrieval and knowledge systems, train and fine-tune models, and deploy AI services in the cloud. You’ll expose these capabilities through secure, documented APIs that our backend developers can integrate into our products.
This role owns AI services and their supporting backend infrastructure. Frontend development is outside its scope.
Responsibilities
- Integrate AI APIs from providers such as OpenAI, Anthropic, and Google, working across different models and their capabilities.
- Build LLM workflows using prompt design, structured outputs, tool calling, streaming, and conversation management.
- Select models based on measured quality, latency, context requirements, and cost; implement routing, fallbacks, and retries where needed.
- Design and maintain retrieval-augmented generation (RAG) systems, including document ingestion, parsing, chunking, embeddings, indexing, retrieval, and reranking.
- Implement and evaluate approaches such as vector RAG, hybrid search, GraphRAG, and agentic retrieval based on product needs.
- Build AI assistants and workflows that connect models with internal data, tools, and business systems.
- Prepare datasets, train models where appropriate, and fine-tune open-source models for specific tasks.
- Deploy and operate open-source models on cloud infrastructure, including GPU environments.
- Develop secure, tested, and documented APIs for backend developers, with clear request formats, responses, and error handling.
- Evaluate model outputs and retrieval quality, and monitor production reliability, latency, and cost.
- Maintain reproducible workflows for experimentation, model versioning, deployment, and ongoing improvement.
- Implement access controls, data protection, and safeguards against prompt injection and unauthorized data retrieval.
Requirements
- Strong Python skills and experience building production backend services with FastAPI or a comparable framework.
- Hands-on experience integrating multiple LLM providers, such as OpenAI, Anthropic, and Google, through their APIs or SDKs.
- Experience building LLM applications beyond basic chat interfaces, including tool calling, structured outputs, and multi-step workflows.
- Practical experience building and operating RAG pipelines, including chunking strategies, embedding selection, metadata filtering, retrieval, and reranking.
- Experience with vector databases or search platforms such as pgvector, Qdrant, Pinecone, Weaviate, or Elasticsearch.
- Understanding of advanced retrieval approaches, with hands-on experience extending basic vector search through hybrid search, GraphRAG, agentic retrieval, or comparable techniques.
- Hands-on experience preparing datasets, training models, and fine-tuning pretrained models, including evaluation and overfitting prevention.
- Experience with PyTorch, Hugging Face, and fine-tuning techniques such as LoRA or QLoRA.
- Proven experience deploying and serving open-source models on AWS, Azure, or Google Cloud using GPU infrastructure.
- Understanding of inference optimization, including quantization, batching, caching, and GPU memory management.
- Experience with API authentication, asynchronous processing, automated testing, Docker, Linux, Git, and CI/CD.
- Ability to measure AI quality and diagnose failures across prompts, retrieval, models, and infrastructure.
Nice to Have
- Experience with frameworks such as LangChain, LangGraph, or LlamaIndex.
- Experience with model-serving tools such as vLLM or NVIDIA Triton.
- Experience with knowledge graphs and graph databases.
- Familiarity with multimodal AI, including document, image, and speech processing.
- Experience with Kubernetes, infrastructure as code, and distributed training or inference.
- Familiarity with experiment tracking, LLM tracing, and automated evaluation tools.
- Experience implementing human review and approval steps in AI workflows.
What Success Looks Like
- AI capabilities meet agreed quality, reliability, latency, and cost targets.
- Retrieval systems return relevant, permission-appropriate information and support grounded answers.
- Backend developers can integrate AI capabilities through clear, stable APIs.
- Hosted APIs and self-hosted models are selected using evidence from evaluations.
- Training, fine-tuning, and deployment workflows are reproducible and maintainable.