AI Jobs Map

Web Spiders · Greater Kolkata Area

Senior AWS EMR Engineer

seniorfull timePosted yesterday
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

awshadoopetls3apache-airflowscalaserverlessredshiftci/cdapache-sparkdata-engineeringartificial-intelligence

Web Spiders is looking for a Senior AWS EMR Engineer with strong hands-on experience in AWS EMR, Apache Spark, Hadoop, and large-scale distributed data processing.

The ideal candidate will have experience building, managing, optimizing, and troubleshooting production-grade data processing workloads on AWS, with a strong understanding of EMR clusters, Spark workloads, data pipelines, performance optimization, scalability, reliability, and cost efficiency.

If AWS EMR + Spark/Hadoop is your core expertise, we'd love to hear from you.

‍5+ Years Experience | Kolkata – Work from Office

Core Stack: AWS EMR

- Apache Spark

- Hadoop

- S3

- Glue

- Airflow/MWAA

- Step Functions

- Immediate joiners preferred.*

Working Hours: Ability to work in the US Eastern Time Zone. Depending on project requirements, this may be adjusted to a half-day IST + half-day US EST schedule.

What You'll Do

- Design, develop, deploy, and optimize large-scale data processing workloads using AWS EMR and Apache Spark.

- Build and maintain distributed data processing solutions using Spark/Hadoop.

- Develop and optimize Spark jobs for performance, scalability, reliability, and cost efficiency.

- Work with PySpark/Scala for distributed data processing and transformation.

- Configure and manage EMR clusters based on workload and processing requirements.

- Optimize Spark applications, including resource utilization, partitioning, joins, caching, and execution performance.

- Troubleshoot EMR, Spark, Hadoop, and production data-processing issues.

- Work with Amazon S3 as a scalable data lake/storage layer.

- Integrate EMR workloads with AWS services such as Glue, Lambda, Step Functions, and Airflow/MWAA.

- Monitor data-processing workloads and implement appropriate logging, error handling, and operational controls.

- Optimize cloud workloads for performance, scalability, reliability, and AWS cost.

- Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data-processing solutions.

Must-Have Skills

- 5+ years of hands-on experience in Data Engineering / Big Data Engineering.

- Strong hands-on experience with AWS EMR.

- Strong experience with Apache Spark and distributed data processing.

- Strong understanding of Hadoop ecosystem and distributed computing concepts.

- Strong programming experience with PySpark and/or Scala.

- Experience working with Amazon S3 and AWS-based data lakes.

- Experience troubleshooting and optimizing Spark/EMR workloads.

- Strong understanding of ETL/ELT concepts and large-scale data processing.

- Experience with production data pipelines and performance optimization.

Good To Have

- AWS Glue

- Apache Airflow / MWAA

- AWS Step Functions

- AWS Lambda

- Amazon Redshift

- Experience with Spark performance tuning and cluster optimization.

- Experience with CI/CD and deployment of data-processing applications.

- AWS Certified Data Engineer or another relevant AWS certification.

Interview Process

- Application review

- 5–10 minute initial screening call with the TA team

- Technical interviews {Domain specific}

- ‍Practical test conducted in the presence of a panel member

- ‍Role match & offer

More jobs at Web Spiders