Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Build and optimize batch and real-time data processing workflows.
- Extract data from databases, APIs, files, applications, and other sources.
- Transform and cleanse large datasets to meet business and analytical requirements.
- Develop data pipelines using Python, SQL, Spark, and relevant data engineering frameworks.
- Design and maintain data warehouses, data lakes, and lakehouse architectures.
- Implement data quality, validation, monitoring, and error-handling processes.
- Optimize data pipelines and queries for performance, scalability, and cost.
- Develop reusable data models for analytics and reporting.
- Work with cloud-based data platforms and services.
- Implement CI/CD and version control practices for data engineering workflows.
- Monitor production pipelines and troubleshoot data and infrastructure issues.
- Collaborate with Data Scientists, BI Developers, Analysts, Software Engineers, and business teams.
- Ensure compliance with data security, governance, privacy, and access-control requirements.
- Document data pipelines, architecture, data models, and operational procedures.
Required Technical Skills
- Strong programming experience with Python.
- Strong proficiency in SQL and relational databases.
- Hands-on experience with ETL/ELT pipelines.
- Experience with Apache Spark / PySpark.
- Experience with data warehousing concepts and dimensional modeling.
- Knowledge of data lake / lakehouse architecture.
- Experience with at least one major cloud platform:
- Microsoft Azure
- AWS
- Google Cloud Platform
- Experience with orchestration tools such as Apache Airflow, Azure Data Factory, or similar.
- Experience with Git and CI/CD practices.
- Strong understanding of data structures, data modeling, and database concepts.
Cloud & Data Technologies
Experience with one or more of the following is preferred:
Azure
- Azure Data Factory
- Azure Data Lake Storage
- Azure Synapse Analytics
- Azure Databricks
- Azure Functions
- Azure Key Vault
AWS
- S3
- Glue
- EMR
- Redshift
- Lambda
- Athena
Data Platforms & Tools
- Databricks
- Snowflake
- Apache Kafka
- Delta Lake
- Power BI / Tableau
- PostgreSQL / SQL Server / MySQL
- NoSQL databases
Good to Have