Senior Data Engineer – Python/PySpark | AWS | Databricks
Role: Senior Data Engineer
Job Summary:
We are seeking a Senior Data Engineer with strong hands-on experience in Python, PySpark, Databricks, and AWS. The candidate will be responsible for developing scalable data pipelines, integrating and transforming structured/unstructured data, ensuring data quality, and supporting production data platforms.
Key Responsibilities:
- Design, develop, and maintain scalable data pipelines and ETL/ELT processes.
- Ingest, integrate, map, cleanse, transform, and validate structured and unstructured data.
- Develop data processing solutions using Python 3, PySpark, and Databricks.
- Build and manage workflows using Apache Airflow.
- Work with AWS cloud services for data engineering solutions.
- Write technical specifications, unit tests, and maintain technical documentation.
- Participate in data architecture and technical design discussions.
- Implement data engineering and data management best practices.
- Develop and maintain CI/CD pipelines using Azure DevOps/TFS/VSTS.
- Use Terraform for infrastructure automation where required.
- Perform API testing using Postman/Insomnia.
- Monitor production pipelines, troubleshoot issues, and provide maintenance/support.
- Ensure timely delivery and adherence to project deadlines.
Required Skills:
- Python 3 – Expert
- PySpark / Apache Spark – Expert
- Databricks – Advanced
- AWS – Advanced
- Apache Airflow
- Azure DevOps / TFS / VSTS
- Terraform
- CI/CD
- Postman / Insomnia
- C# – Good to have