Job description:
Job Title: AI Engineer (LLM & GenAI)
Experience: 5 to 8 years
Employment Type: Full-Time
Our client is seeking an ambitious AI Engineer with hands-on experience in building, deploying, and evaluating LLM-powered applications to join our growing engineering team.
Role Overview
As a AI Engineer, you will bridge the gap between cutting-edge GenAI research and robust, production-ready software. You will be responsible for building end-to-end LLM solutions from prompt architecture and Retrieval-Augmented Generation (RAG) pipelines to fine-tuning, latency optimization, and agentic workflows.
You will work closely with product managers, data engineers, and full-stack developers to ship high-impact features directly to production.
Key Responsibilities
- LLM Architecture & Integration: Design and implement production-grade LLM applications using commercial APIs (OpenAI, Anthropic) and open-source foundation models (Llama, Mistral).
- RAG & Search Systems: Build and optimize end-to-end Retrieval-Augmented Generation pipelines using vector databases (e.g., Pinecone, Weaviate, pgvector), chunking strategies, and hybrid search mechanics.
- Agentic Workflows: Orchestrate multi-step LLM agents, function calling, and tool-use frameworks (e.g., LangGraph, AutoGen) to solve complex, non-linear workflows.
- Model Optimization & Fine-Tuning: Apply techniques like PEFT/LoRA and quantization to fine-tune open-source models for domain-specific tasks when necessary.
- Evaluation & Guardrails: Establish robust AI evaluation suites (latency, hallucination detection, cost tracking, accuracy) using frameworks like Ragas or TruLens, ensuring safe and reliable AI output.
- API & Infrastructure Integration: Package AI workflows into clean REST or gRPC microservices (FastAPI/Python) and deploy them via Docker/Kubernetes in modern cloud environments.
Qualifications & RequirementsMust-Haves:
- Experience: 5 to 8 years of experience in software engineering, machine learning, or data engineering, with at least 2–3 years directly dedicated to building with LLMs/Generative AI.
- Core Language: Advanced proficiency in Python (PyTorch, NumPy, Pandas) and async backend programming (FastAPI, Flask).
- LLM Tooling: Hands-on experience with LLM orchestration frameworks (LangChain, LlamaIndex) and provider APIs.
- Vector DBs: Practical experience working with vector search engines (Pinecone, Qdrant, Chroma, or pgvector).
- Software Fundamentals: Strong understanding of REST APIs, Git workflows, CI/CD pipelines, and microservices architecture.
- Work Authorization: Must be legally authorized to work in the United States.
Nice-to-Have:
- Bachelor’s or Master’s degree in Computer Science, Data Science, or a related quantitative field.
- Experience with agentic frameworks (LangGraph, CrewAI).
- Familiarity with cloud platforms (AWS Bedrock/SageMaker, GCP Vertex AI, Azure OpenAI).
- Experience handling LLM observability and monitoring tools (LangSmith, Phoenix, Arize).
Pay: $150,000.00 - $170,000.00 per year
Work Location: On the road