Experience: 3+ Years
Employment Type: Full-Time
Work Mode: Remote
Role Type: Cloud Data Modernization
Job Summary
We are seeking a highly skilled Cloud Data Engineer with strong hands-on experience in Databricks, Azure Cloud, Snowflake, Apache Spark, PySpark, and ETL development to join our team for a large-scale Cloud Data Modernization and Migration project.
The ideal candidate will have experience migrating on-premises ETL workloads to modern cloud data platforms, developing scalable data pipelines, implementing CI/CD and DevOps practices, and supporting data quality, reconciliation, monitoring, and observability.
Databricks is the primary data engineering skill, with Snowflake as a secondary skill. Experience with Azure and/or AWS cloud infrastructure is highly desirable.
Primary Skills
- Databricks
- Apache Spark
- PySpark
- Azure Data Factory (ADF)
- Azure Cloud Data Engineering
- Snowflake
- ETL / ELT Development
- Python
- SQL
- Data Migration & Cloud Modernization
- Azure Data Lake Storage Gen2 (ADLS Gen2)
Secondary Skills
- Self-hosted Integration Runtime (SHIR)
- Azure Logic Apps
- Azure Blob Storage
- Delta Lake
- Lakehouse Architecture
- Medallion Architecture – Bronze / Silver / Gold
- GitHub Actions
- CI/CD & DevOps
- Data Quality & Validation
- Data Reconciliation
- Data Observability & Monitoring
- Data Modeling & Database Design
- Apache Airflow
- Snowflake Development
- Performance Optimization
- AI-Assisted Development
- Data Governance
Key Responsibilities
- Contribute to the migration of on-premises ETL workloads to Azure, Databricks, and Snowflake.
- Design, develop, test, and maintain scalable data pipelines using Databricks, PySpark, Apache Spark, ADF, ADLS Gen2, Logic Apps, Blob Storage, and Snowflake.
- Review and analyze existing on-premises ETL processes and identify opportunities for modernization and optimization.
- Develop reliable ETL/ELT pipelines and data processing workflows using Python, PySpark, SQL, and cloud data services.
- Implement and maintain CI/CD pipelines using GitHub Actions and follow DevOps engineering practices.
- Work with GitHub branching strategies, pull requests, code reviews, and version control.
- Collaborate with data architects, analysts, developers, QA teams, and business stakeholders to ensure seamless data integration and flow.
- Tune and optimize Databricks and Snowflake workloads for performance, scalability, and cost efficiency.
- Implement data quality validation, reconciliation, monitoring, and observability processes.
- Develop and maintain automated data reconciliation and validation frameworks.
- Support automated testing and troubleshoot data pipeline failures and data-quality issues.
- Ensure appropriate data security, governance, and compliance practices are followed.
- Participate in Agile ceremonies and contribute to continuous improvement of engineering processes.
- Work independently, manage ambiguity, identify critical issues, and deliver high-quality solutions within established timelines.
- Leverage AI-assisted development tools such as GitHub Copilot, Databricks Assistant, ChatGPT, or Claude to improve engineering productivity.
Required Qualifications
- 3+ years of experience as a Cloud Data Engineer or Data Engineer.
- Strong hands-on experience with Databricks, Apache Spark, PySpark, and Python.
- Experience with Azure Data Factory (ADF), ADLS Gen2, Blob Storage, SHIR, and/or Logic Apps.
- Strong experience with Snowflake for data engineering and analytics workloads.
- Hands-on experience developing and supporting ETL/ELT pipelines.
- Experience with on-premises databases and ETL technologies and migration to cloud platforms.
- Strong SQL development skills, including:
- Complex SQL queries
- Stored procedures
- Views
- Joins and transformations
- Data validation and reconciliation
- Experience with GitHub, branching, pull requests, and GitHub Actions.
- Experience with code reviews and software/data engineering best practices.
- Experience supporting automated testing, data quality, and data validation processes.
- Experience working in Agile/Scrum delivery environments.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to learn new technologies quickly and work effectively in a fast-paced environment.
- Experience using AI-assisted development tools such as GitHub Copilot, Databricks Assistant, ChatGPT, Claude, or similar tools.
Preferred Qualifications
- Experience with AWS or other cloud platforms in addition to Azure.
- Experience implementing Bronze/Silver/Gold (Medallion) architecture.
- Strong knowledge of Delta Lake and Lakehouse architecture.
- Experience with Delta Lake optimization techniques.
- Experience with Snowflake data engineering and analytics development.
- Experience with data modeling and database design.
- Knowledge of data governance and data quality best practices.
- Experience with Apache Airflow or other orchestration frameworks.
- Experience optimizing workloads in Databricks and Snowflake.
- Experience leveraging AI/ML solutions for data engineering workflows and automation.
- Experience working with healthcare payer data, including:
- Claims
- Membership
- Enrollment
- Provider
- Clinical
- Financial data
- Azure or Databricks certifications.
Cloud & Technology Environment
Primary:
Databricks, PySpark, Apache Spark, Azure Data Factory, ADLS Gen2, Snowflake
Cloud:
Azure and/or AWS
Data Engineering:
ETL/ELT, Python, SQL, Delta Lake, Lakehouse, Medallion Architecture
DevOps:
GitHub, GitHub Actions, CI/CD, Code Reviews
Data Quality:
Data Validation, Reconciliation, Testing, Monitoring, Observability
Orchestration:
ADF, Logic Apps, Apache Airflow
AI-Assisted Engineering:
GitHub Copilot, Databricks Assistant, ChatGPT, Claude
Work Location: Remote