Data Engineer – Databricks | Pharma Domain
Location: India (Remote)
Experience: 4+ Years
Employment Type: Contract / Full-Time
Domain: Pharma Domain – Mandatory
Job Summary
We are looking for an experienced Data Engineer with strong hands-on experience in Databricks, PySpark, SQL, CI/CD, and Version Control. The ideal candidate should have experience working on Pharma domain projects and be comfortable developing and maintaining scalable data pipelines.
Key Responsibilities
- Design, develop, and maintain data pipelines using Databricks.
- Work with Databricks Notebooks and Databricks Pipelines for data processing and transformation.
- Develop data transformation and processing workflows using PySpark.
- Write optimized SQL queries for data extraction, transformation, and analysis.
- Build and maintain reliable ETL/ELT pipelines for large datasets.
- Implement data quality, validation, and error-handling processes.
- Work with CI/CD pipelines for deployment and release management.
- Use Git/version control for source-code management and collaborative development.
- Collaborate with business and technical teams to understand data requirements.
- Follow security, governance, documentation, and data engineering best practices.
Mandatory Skills
- Pharma Domain Experience – MUST
- Databricks – MUST
- Databricks Notebooks
- Databricks Pipelines
- PySpark – Strong/basic hands-on experience
- SQL – Strong
- CI/CD
- Git / Version Control
- Strong understanding of ETL/ELT and data pipeline development
Preferred
- Experience with cloud platforms such as Azure / AWS
- Experience working with large-scale datasets
- Understanding of data quality and data governance
- Experience working in Agile development environments
Important
Candidates must have direct Pharma domain/project experience.
Healthcare or non-Pharma experience will not be considered.