AI Jobs Map

Web Spiders · Greater Kolkata Area

Senior Spark / PySpark Data Engineer

seniorfull timePosted yesterday
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

apache-sparkdata-engineeringpythonetlawss3redshiftartificial-intelligenceapache-airflowhadoopdata-warehousingci/cd

Web Spiders is looking for a Senior Spark / PySpark Data Engineer with strong hands-on experience in Apache Spark, PySpark, Python, and large-scale data engineering.

The ideal candidate will have strong experience designing and developing high-performance ETL/ELT pipelines and distributed data-processing solutions using Spark/PySpark, along with experience working with cloud-based data platforms and AWS services.

If Spark + PySpark + Python is your core expertise and you enjoy solving complex data-processing and scalability challenges, we'd love to hear from you.

5+ Years Experience | Kolkata – Work from Office

Core Stack: Apache Spark

- PySpark

- Python

- ETL/ELT

- AWS

- S3

- Glue

- Redshift

- Immediate joiners preferred.*

Working Hours: Ability to work in the US Eastern Time Zone. Depending on project requirements, this may be adjusted to a half-day IST + half-day US EST schedule.

What You'll Do

- Design, develop, and optimize large-scale ETL/ELT pipelines using Apache Spark and PySpark.

- Develop scalable data transformation and processing solutions using PySpark and Python.

- Build distributed data-processing applications capable of handling large volumes of data.

- Develop reusable and maintainable Spark/PySpark frameworks and data-processing components.

- Optimize Spark jobs for performance, scalability, memory utilization, and execution efficiency.

- Work with complex transformations, joins, aggregations, partitioning, and large datasets.

- Implement data validation, quality checks, error handling, and monitoring within data pipelines.

- Work with AWS data services including EMR, Glue, S3, and Redshift.

- Develop data pipelines supporting data lakes, warehouses, analytics, and downstream applications.

- Troubleshoot production data pipeline and Spark processing issues.

- Identify and resolve performance bottlenecks in Spark/PySpark workloads.

- Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.

Must-Have Skills

- 5+ years of hands-on experience in Data Engineering.

- Strong hands-on experience with Apache Spark.

- Strong hands-on experience with PySpark.

- Strong programming experience in Python.

- Proven experience developing and optimizing large-scale ETL/ELT pipelines.

- Strong understanding of distributed computing and data-processing concepts.

- Experience working with large datasets and complex data transformations.

- Strong understanding of Spark performance optimization and tuning.

- Experience with cloud-based data engineering, preferably AWS.

- Experience with Amazon S3 and at least one AWS data-processing service such as EMR or Glue.

Good To Have

- AWS EMR

- AWS Glue

- Apache Airflow / MWAA

- AWS Step Functions

- Amazon Redshift

- Hadoop ecosystem

- Experience with data lake and data warehouse architectures.

- Experience with CI/CD and production deployment of data pipelines.

- AWS Certified Data Engineer or another relevant AWS certification.

Interview Process

- Application review

- 5–10 minute initial screening call with the TA team

- Technical interviews {Domain specific}

- ‍Practical test conducted in the presence of a panel member

- ‍Role match & offer

More jobs at Web Spiders