About AGS Health
AGS Health is a leading strategic growth partner to healthcare providers across the U.S. Working alongside each client as one team, we improve revenue cycle performance by orchestrating data, AI, automation, and human expertise across every workflow. Our proven model complements customers' existing teams, technology, and processes, deploying the right capabilities to improve efficiency, strengthen financial performance, and enhance the patient financial experience. Supported by a global team of more than 16,000 clinical and revenue cycle experts across onshore, nearshore, and offshore locations, AGS Health helps customers achieve measurable results across diverse care settings and specialties.
About the Role
Build the data pipelines that make production machine learning and LLM systems possible: curated training/evaluation datasets, a labeling pipeline with real quality control, and production model monitoring for a healthcare claims and coding workload.
Key Responsibilities
- Build pipelines that extract, curate, and structure training and evaluation datasets from historical claims/coding records for use in machine learning and LLM systems.
- Own labeling-pipeline design and quality control for supervised or semi-supervised model training.
- Stand up production model monitoring: detect data and prediction drift once a model is live, and alert on out-of-tolerance shifts.
- Provide clean, representative, correctly time-ordered training and evaluation data (train-old/test-new discipline, avoiding data leakage) to ML/AI engineering.
- Partner with clinical/coding subject-matter experts to validate that labeled data reflects real-world edge cases in claims adjudication or medical coding.
- Write production-grade SQL and Python for data extraction, transformation, and pipeline automation.
Required Qualifications
- Production-level SQL and Python - not analyst-level scripting - with experience building and maintaining automated data pipelines.
- Direct experience with healthcare claims data (medical, dental, or pharmacy) including standard coding systems (ICD-10-CM, CPT/HCPCS).
- Experience with data-quality controls for a labeling/annotation pipeline feeding a machine learning model.
Preferred Qualifications
- Experience building ML model monitoring or drift detection in a production system.
- Exposure to vector embeddings or retrieval pipelines for LLM-based systems.
- Statistical or ML modeling background (scikit-learn, XGBoost, or similar).
- Technology & Tools
- Languages: Python (pandas, NumPy, scikit-learn) and SQL at a production standard.
- Data warehouse/lake: Snowflake, Redshift, BigQuery, or AWS S3/Glue/EMR.
- Pipeline orchestration: Airflow, dbt, or an equivalent scheduling/transformation framework.
- Labeling tooling: Label Studio, Snorkel, or an equivalent annotation platform with inter-rater quality tracking.
- Model monitoring: Evidently, Arize, WhyLabs, or an equivalent drift-detection tool.
- Exposure to embeddings/vector stores (pgvector, OpenSearch) a plus for supporting retrieval-based systems.
Education
Bachelor's degree in Computer Science, Statistics, Data Science, or a related quantitative field. Master's degree a plus.
Certifications
No certification required. A healthcare data/health-information credential (e.g., CHDA) is a plus.
Relevant Background
Direct experience at a healthcare claims/analytics company, a payer or RCM vendor with a claims data warehouse, or a data engineering team within a health-tech company supporting a live ML/LLM product - not a purely academic or general-BI analytics