Job Description
Job Description – GCP Data EngineerRole Overview
We are looking for an experienced GCP Data Engineer to design, develop, and maintain scalable data pipelines and ETL/ELT frameworks on the Google Cloud Platform. The candidate will work extensively with BigQuery, Python/PySpark, Airflow/Cloud Composer, and Google Cloud Storage to process and transform large volumes of campaign, customer, and clickstream data.
The ideal candidate should have strong hands-on experience in data engineering, excellent SQL and Python skills, and a good understanding of cloud-based data pipeline architecture. The role will involve building reliable production pipelines, optimizing BigQuery performance and costs, implementing automation, and supporting critical data workflows.
Key ResponsibilitiesData Pipeline Development
- Design, develop, and maintain scalable and reliable data pipelines on GCP.
- Build batch and, where required, near-real-time data ingestion and transformation pipelines.
- Develop robust ETL/ELT frameworks for large volumes of structured and semi-structured data.
- Ingest data from multiple sources into Google Cloud Storage and BigQuery.
- Develop reusable and modular data engineering components.
- Implement appropriate error handling, retry mechanisms, logging, and data validation.
BigQuery & Data Engineering
- Develop complex and optimized SQL queries in BigQuery for data transformation and analysis.
- Optimize BigQuery performance through:
- Partitioning
- Clustering
- Query optimization
- Efficient table design
- Appropriate data types and storage strategies
- Monitor and optimize BigQuery processing costs.
- Design scalable data models suitable for large-scale campaign, customer, and clickstream datasets.
- Troubleshoot data quality, performance, and pipeline-related issues.
Airflow / Cloud Composer
- Develop and maintain DAGs using Apache Airflow / Cloud Composer.
- Implement scheduling, dependency management, retries, failure handling, and alerting.
- Monitor production workflows and proactively resolve failed or delayed jobs.
- Build reusable operators and workflow components where required.
- Ensure critical data pipelines meet agreed SLAs.
Python / PySpark
- Develop data processing and transformation logic using Python and/or PySpark.
- Write clean, reusable, scalable, and production-ready code.
- Optimize PySpark jobs for performance and efficient resource utilization.
- Implement appropriate testing and validation for data transformation processes.
GCP Services
Work extensively with GCP services, particularly:
- Google BigQuery
- Google Cloud Storage (GCS)
- Cloud Composer / Airflow
Exposure to additional GCP services such as Cloud Functions, Pub/Sub, Dataflow, Cloud Run, Secret Manager, IAM, or Cloud Monitoring would be an advantage.
Automation, CI/CD & DevOps
- Automate deployment and execution of data engineering workflows.
- Work with Git-based development and version-control workflows.
- Implement and maintain CI/CD pipelines using tools such as Jenkins or equivalent.
- Follow code review, branching, deployment, and release-management processes.
- Implement monitoring, logging, and alerting for production data pipelines.
Production Support & Troubleshooting
- Monitor daily data pipeline execution and resolve production issues within defined SLAs.
- Investigate pipeline failures, data discrepancies, performance issues, and processing delays.
- Perform root-cause analysis and implement permanent fixes.
- Coordinate with application, analytics, infrastructure, and business teams to resolve data-related issues.
- Participate in production deployments and provide post-deployment support.
Stakeholder Collaboration
- Work closely with Data Analysts, Data Scientists, Product Teams, Business Stakeholders, and Technology Teams to understand data requirements.
- Translate business requirements into scalable technical solutions.
- Communicate technical issues, risks, dependencies, and delivery status effectively.
- Participate in technical discussions, design reviews, and solution development.
Required Technical SkillsMust Have
- 4–8 years of experience in Data Engineering.
- Strong hands-on experience with GCP.
- Strong proficiency in SQL, preferably extensive experience with BigQuery SQL.
- Strong hands-on experience in Python and/or PySpark.
- Experience developing ETL/ELT data pipelines.
- Hands-on experience with BigQuery.
- Hands-on experience with Google Cloud Storage (GCS).
- Experience with Apache Airflow / Cloud Composer.
- Good understanding of data pipeline architecture and data engineering best practices.
- Experience with Git and CI/CD practices.
- Exposure to Jenkins or similar CI/CD tools.
- Experience in production support, monitoring, troubleshooting, and performance optimization.
Good to Have
- Experience working with campaign, marketing, customer, or clickstream data.
- Experience with large-scale data processing.
- Knowledge of Dataflow / Apache Beam.
- Knowledge of Pub/Sub and event-driven architectures.
- Experience with real-time or streaming data pipelines.
- Knowledge of data warehousing and dimensional data modeling.
- Experience with data quality frameworks and validation.
- Knowledge of GCP IAM and security concepts.
- Experience with Cloud Monitoring / Logging.
- Experience working in Agile/Scrum environments.
Candidate Profile
The ideal candidate should:
- Have strong hands-on technical expertise rather than only theoretical knowledge.
- Be comfortable writing complex SQL and Python/PySpark code.
- Have experience independently designing and developing data pipelines.
- Understand how to build scalable and cost-efficient solutions on GCP.
- Be capable of troubleshooting production issues and taking ownership until resolution.
- Have good analytical and problem-solving skills.
- Be comfortable working with multiple stakeholders and managing delivery timelines.
- Demonstrate good communication and documentation skills.
About The Team
eClerx is a global leader in productized services, bringing together people, technology and domain expertise to amplify business results. Our mission is to set the benchmark for client service and success in our industry. Our vision is to be the innovation partner of choice for technology, data analytics and process management services. Since our inception in 2000, we've partnered with top companies across various industries, including financial services, telecommunications, retail, and high-tech. Our innovative solutions and domain expertise help businesses optimize operations, improve efficiency, and drive growth. With over 18,000 employees worldwide, eClerx is dedicated to delivering excellence through smart automation and data-driven insights. At eClerx, we believe in nurturing talent and providing hands-on experience.
eClerx is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability or protected veteran status, or any other legally protected basis, per applicable law.