PLEASE NOTE BEFORE APPLYING:
CODOXO IS NOT ABLE TO OFFER SPONSORSHIP OR ACCOMMODATE ANY CANDIDATES THAT ARE CURRENTLY BEING SPONSORED NOW OR IN THE FUTURE
The United States spends roughly $4.9 trillion on healthcare each year, and an estimated quarter of that is lost to waste, fraud, abuse, and error. Codoxo is the premier provider of AI-driven solutions that help healthcare companies and government agencies proactively detect and reduce those losses and ensure payment integrity.
We are purpose-driven, with the goal of making healthcare more affordable and accessible to all. If you are passionate about applied AI and driven by positive impact in healthcare, Codoxo is where you belong. We are venture backed by some of the top investors in the country, with strong financials, and remain one of the fastest growing healthcare AI companies in the industry.
Position Summary
As an AI/LLM Engineer, you will design, ship, and operate LLM-powered features that accelerate claim audits and SIU investigations — retrieval and question answering over case documents, structured extraction from claims and medical records, and investigator-facing summarization. Just as importantly, you will build the evaluation systems that tell us whether any of it actually works.
Our production stack is Python and Django on AWS, with Amazon Bedrock for inference, OpenSearch for vector and hybrid search, and Celery for asynchronous document processing. You will work on a small team where you own features end to end and your work reaches fraud investigators at national health plans and government agencies.
As our LLM foundation matures, this role grows into multi-step, tool-using systems — MCP tools, Bedrock AgentCore, and multi-agent orchestration. You would help decide when that complexity is earned rather than inheriting the decision. We are looking for someone who reaches for the simplest thing that clears the bar and can explain why.
Key Responsibilities
-
Build and ship LLM features against real investigator workflows: document question answering, structured extraction, and summarization over claims, medical records, notes, and correspondence.
-
Own our retrieval pipeline end to end — document ingestion and extraction quality, chunking, metadata filtering, hybrid lexical and semantic search, reranking, and citations traced back to source text.
-
Build and maintain the evaluation systems that gate our releases: curated golden sets built with subject-matter experts, regression suites that run on every prompt and model change, and rubric-based grading validated against human reviewers.
-
Treat prompts as engineering artifacts — versioned, code-reviewed, documented with change notes, and never shipped without an evaluation run.
-
Design reliable structured output: schema-constrained generation, validation with bounded retries, and graceful handling of refusals, truncation, and partial results.
-
Build on serverless AWS — Bedrock, Lambda, S3, API Gateway, Step Functions — keeping PHI inside our trust boundary and out of logs and traces.
-
Keep cost and latency predictable as usage grows: token accounting, prompt caching, routing tasks to the right model tier, and asynchronous processing for non-interactive work.
-
Instrument LLM features for observability and audit — prompt and model versions, token counts, validation and retry rates, retrieved document IDs.
-
Partner with investigators, clinical coders, and product to turn expert judgment into evaluation criteria, not just requirements.
Qualifications
-
3+ years of professional Python experience.
-
2+ years building with LLMs in production — not prototypes or demos. You have shipped something real, watched it fail in ways you did not expect, and fixed it.
-
Hands-on RAG experience: embeddings, vector or hybrid search (OpenSearch, pgvector, FAISS, or similar), chunking strategy, and the judgment to diagnose whether a bad answer came from retrieval or generation.
-
A real practice around evaluation. You can describe how you knew a feature was working, what your eval set looked like, and a time a metric misled you.
-
Working knowledge of AWS serverless — Lambda, S3, API Gateway — and at least one managed LLM service. Bedrock experience is a plus, but strong OpenAI, Azure OpenAI, or Vertex AI experience transfers fine.
-
Comfort working with PHI or other regulated data, or genuine interest in learning to do it properly.
If you meet most of this list, we would rather see your application than not. We are more interested in how you reason about LLM systems than in whether you have used our exact stack.
Bonus Points
-
Healthcare domain data: FHIR, ICD-10, CPT, HCPCS, and payer claims experience. Valuable, and something we can teach the right engineer.
-
Document extraction and OCR pipelines — the unglamorous work that determines whether everything downstream succeeds.
-
Amazon Bedrock AgentCore (Memory, Gateway, Runtime), Strands Agents, MCP servers, or A2A.
-
Event-driven AWS: Step Functions, EventBridge, Glue/Athena.
-
Observability practice: CloudWatch, OpenTelemetry, structured logging, incident response.
-
Parameter-Efficient Fine-Tuning (LoRA/QLoRA) and knowing when it is the wrong tool.
-
Agent frameworks: LangGraph, LlamaIndex, CrewAI, or equivalent.
-
Kubernetes/EKS, GPU inference, and caching strategy.
-
Explainability, bias monitoring, and responsible-AI controls.
Physical Requirements
Work is performed in an office environment (either in our office or work-from-home) and requires the ability to work on a computer, operate standard office equipment, and work at a desk.
Benefits for You
-
Competitive equity.
-
Health, dental, and vision insurance with 100% employee premium coverage (starts day 1).
-
Unlimited PTO.
-
Annual professional development stipend.
-
Annual home office stipend.
-
401K match (after 90 days).
Accessibility Notice
If you need reasonable accommodation for any part of the employment process due to a physical or mental disability, please send an email to [email protected] with the subject “Accommodation”. Reasonable accommodation requests will be considered on a case-by-case basis.
We Are an Equal Opportunity Employer
Codoxo prohibits discrimination of any type and affords equal employment opportunities to employees and applicants without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by law. This policy applies to all terms and conditions of employment.