Web Spiders is looking for a Senior AWS EMR Engineer with strong hands-on experience in AWS EMR, Apache Spark, Hadoop, and large-scale distributed data processing. The ideal candidate will have experience building, managing, optimizing, and troubleshooting production-grade data processing workloads on AWS, with a strong understanding of EMR clusters, Spark workloads, data pipelines, performance optimization, scalability, reliability, and cost efficiency.
Kolkata (Rajarhat–Newtown)
Web Spiders is looking for a Senior AWS EMR Engineer with strong hands-on experience in AWS EMR, Apache Spark, Hadoop, and large-scale distributed data processing.
The ideal candidate will have experience building, managing, optimizing, and troubleshooting production-grade data processing workloads on AWS, with a strong understanding of EMR clusters, Spark workloads, data pipelines, performance optimization, scalability, reliability, and cost efficiency.
If AWS EMR + Spark/Hadoop is your core expertise, we'd love to hear from you.
5+ Years Experience | Kolkata – Work from Office
Core Stack: AWS EMR • Apache Spark • Hadoop • S3 • Glue • Airflow/MWAA • Step Functions
* Immediate joiners preferred.*
Working Hours: Ability to work in the US Eastern Time Zone. Depending on project requirements, this may be adjusted to a half-day IST + half-day US EST schedule.
What You'll Do:
Design, develop, deploy, and optimize large-scale data processing workloads using AWS EMR and Apache Spark.
Build and maintain distributed data processing solutions using Spark/Hadoop.
Develop and optimize Spark jobs for performance, scalability, reliability, and cost efficiency.
Work with PySpark/Scala for distributed data processing and transformation.
Configure and manage EMR clusters based on workload and processing requirements.
Optimize Spark applications, including resource utilization, partitioning, joins, caching, and execution performance.
Troubleshoot EMR, Spark, Hadoop, and production data-processing issues.
Work with Amazon S3 as a scalable data lake/storage layer.
Integrate EMR workloads with AWS services such as Glue, Lambda, Step Functions, and Airflow/MWAA.
Monitor data-processing workloads and implement appropriate logging, error handling, and operational controls.
Optimize cloud workloads for performance, scalability, reliability, and AWS cost.
Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data-processing solutions.
Must-Have Skills:
5+ years of hands-on experience in Data Engineering / Big Data Engineering.
Strong hands-on experience with AWS EMR.
Strong experience with Apache Spark and distributed data processing.
Strong understanding of Hadoop ecosystem and distributed computing concepts.
Strong programming experience with PySpark and/or Scala.
Experience working with Amazon S3 and AWS-based data lakes.
Experience troubleshooting and optimizing Spark/EMR workloads.
Strong understanding of ETL/ELT concepts and large-scale data processing.
Experience with production data pipelines and performance optimization.
Good to Have:
AWS Glue
Apache Airflow / MWAA
AWS Step Functions
AWS Lambda
Amazon Redshift
Experience with Spark performance tuning and cluster optimization.
Experience with CI/CD and deployment of data-processing applications.
AWS Certified Data Engineer or another relevant AWS certification.
Interview Process
Application review
5–10 minute initial screening call with the TA team
Technical interviews {Domain specific}
Practical test conducted in the presence of a panel member
Role match & offer