AI Jobs Map

Flipped.ai · New York, NY

AI Engineer- LLM and Gen AI

Hybridfull time$150,000 – $170,000 / yearPosted today
Apply on IndeedIndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmgenerative-airagagentic-aiopenaianthropicvector-databasespineconeweaviatefine-tuninggrpcmicroservicesfastapipythondockerkubernetesmachine-learningdata-engineeringpytorchnumpy

Job description:

Job Title: AI Engineer (LLM & GenAI)

Experience: 5 to 8 years

Employment Type: Full-Time

Our client is seeking an ambitious AI Engineer with hands-on experience in building, deploying, and evaluating LLM-powered applications to join our growing engineering team.

Role Overview

As a AI Engineer, you will bridge the gap between cutting-edge GenAI research and robust, production-ready software. You will be responsible for building end-to-end LLM solutions from prompt architecture and Retrieval-Augmented Generation (RAG) pipelines to fine-tuning, latency optimization, and agentic workflows.

You will work closely with product managers, data engineers, and full-stack developers to ship high-impact features directly to production.

Key Responsibilities

- LLM Architecture & Integration: Design and implement production-grade LLM applications using commercial APIs (OpenAI, Anthropic) and open-source foundation models (Llama, Mistral).

- RAG & Search Systems: Build and optimize end-to-end Retrieval-Augmented Generation pipelines using vector databases (e.g., Pinecone, Weaviate, pgvector), chunking strategies, and hybrid search mechanics.

- Agentic Workflows: Orchestrate multi-step LLM agents, function calling, and tool-use frameworks (e.g., LangGraph, AutoGen) to solve complex, non-linear workflows.

- Model Optimization & Fine-Tuning: Apply techniques like PEFT/LoRA and quantization to fine-tune open-source models for domain-specific tasks when necessary.

- Evaluation & Guardrails: Establish robust AI evaluation suites (latency, hallucination detection, cost tracking, accuracy) using frameworks like Ragas or TruLens, ensuring safe and reliable AI output.

- API & Infrastructure Integration: Package AI workflows into clean REST or gRPC microservices (FastAPI/Python) and deploy them via Docker/Kubernetes in modern cloud environments.

Qualifications & RequirementsMust-Haves:

- Experience: 5 to 8 years of experience in software engineering, machine learning, or data engineering, with at least 2–3 years directly dedicated to building with LLMs/Generative AI.

- Core Language: Advanced proficiency in Python (PyTorch, NumPy, Pandas) and async backend programming (FastAPI, Flask).

- LLM Tooling: Hands-on experience with LLM orchestration frameworks (LangChain, LlamaIndex) and provider APIs.

- Vector DBs: Practical experience working with vector search engines (Pinecone, Qdrant, Chroma, or pgvector).

- Software Fundamentals: Strong understanding of REST APIs, Git workflows, CI/CD pipelines, and microservices architecture.

- Work Authorization: Must be legally authorized to work in the United States.

Nice-to-Have:

- Bachelor’s or Master’s degree in Computer Science, Data Science, or a related quantitative field.

- Experience with agentic frameworks (LangGraph, CrewAI).

- Familiarity with cloud platforms (AWS Bedrock/SageMaker, GCP Vertex AI, Azure OpenAI).

- Experience handling LLM observability and monitoring tools (LangSmith, Phoenix, Arize).

Pay: $150,000.00 - $170,000.00 per year

Work Location: On the road

More jobs at Flipped.ai

Similar roles in New York