Skills
Azure DataBricks ADF, DataBricks, Python, Pyspark, ETL, SQL, LLM, RAG, AI/ML.
Mandatory skills
- 10+ years in data engineering, with recent hands-on delivery ownership.
- Strong Databricks and PySpark jobs, workflows, Delta Lake, performance tuning
- Expert SQL including the ability to read unfamiliar, poorly documented SQL at volume and reason about what it does and what it costs
- Demonstrable query cost and performance optimization experience, this is a core requirement on this engagement, not a bonus
- Incremental / CDC ingestion pattern** design and implementation - Azure data services — ADLS / Blob Storage, and Databricks on Azure.
- Unity Catalog or comparable data governance and cataloguing experience - Data profiling and reconciliation testing proving two pipelines produce equivalent output
- Comfort working with PHI-scoped healthcare data and the access controls that implies -
- Ability to work independently against ambiguous inputs and drive clarification, in a small team on a hard deadline