Job
Location: 1001 Sagittarius Avenue, Franklin, IN 46131
Job Duties:
- Create and enhance end-to-end data engineering solutions on Azure using Databricks, PySpark, Python, and SQL, handling the ingestion and transformation of high-volume structured and semi-structured data from multiple source systems.
- Develop automated batch and near real-time processing workflows using Azure Data Factory, Databricks notebooks and workflows, Spark, and Delta Lake, with Bronze, Silver, and Gold data layers to organize and process data efficiently.
- Build data solutions based on enterprise architecture and business requirements, including data modeling, schema development, data quality validation, and transformation logic required for reporting, analytics, and downstream applications.
- Implement ingestion processes to bring data from internal and external applications into the Azure data platform, including SAP integrations through MuleSoft, APIs, files, and other interfaces, and perform required data standardization and transformation.
- Develop and maintain Azure Data Lake and Databricks Lakehouse solutions using ADLS and Delta Lake, while using Unity Catalog for data governance, catalog management, lineage, permissions, and controlled access to enterprise data assets.
- Write and maintain data processing logic in Databricks using PySpark, Python, Spark SQL, and Delta tables, and improve workload performance through partitioning, cluster configuration, caching, query tuning, and optimization of Spark jobs.
- Build integrated data pipelines using Azure Data Factory, Azure Databricks, ADLS, Event Hubs, and Key Vault to support data ingestion, orchestration, transformation, storage, and both batch and streaming processing requirements.
- Apply data security and governance standards across Azure and Databricks environments using Unity Catalog, Azure Key Vault, role-based access controls, encryption, tokenization, and secrets management to protect sensitive healthcare, PHI, and PII data.
- Develop monitoring and operational controls for data pipelines using Azure Monitor, Log Analytics, Databricks monitoring, and Python or PySpark utilities, including automated logging, failure notifications, data validation, error handling, and pipeline performance tracking.
- Contribute to data engineering design discussions, code reviews, deployments, troubleshooting, performance tuning, and Agile development activities, while supporting CI/CD practices, reusable engineering standards, and continuous improvements to the Azure Databricks data platform
All the responsibilities mentioned above are in line with the professional background and requires an absolute minimum of a Bachelor’s degree in computer science, computer information systems, information technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor’s degree in one of the aforementioned subjects.