Web Spiders is looking for a Senior Spark / PySpark Data Engineer with strong hands-on experience in Apache Spark, PySpark, Python, and large-scale data engineering. The ideal candidate will have strong experience designing and developing high-performance ETL/ELT pipelines and distributed data-processing solutions using Spark/PySpark, along with experience working with cloud-based data platforms and AWS services.
Kolkata (Rajarhat–Newtown)
Web Spiders is looking for a Senior Spark / PySpark Data Engineer with strong hands-on experience in Apache Spark, PySpark, Python, and large-scale data engineering.
The ideal candidate will have strong experience designing and developing high-performance ETL/ELT pipelines and distributed data-processing solutions using Spark/PySpark, along with experience working with cloud-based data platforms and AWS services.
If Spark + PySpark + Python is your core expertise and you enjoy solving complex data-processing and scalability challenges, we'd love to hear from you.
5+ Years Experience | Kolkata – Work from Office
Core Stack: Apache Spark • PySpark • Python • ETL/ELT • AWS • S3 • Glue • Redshift
* Immediate joiners preferred.*
Working Hours: Ability to work in the US Eastern Time Zone. Depending on project requirements, this may be adjusted to a half-day IST + half-day US EST schedule.
What You'll Do:
Design, develop, and optimize large-scale ETL/ELT pipelines using Apache Spark and PySpark.
Develop scalable data transformation and processing solutions using PySpark and Python.
Build distributed data-processing applications capable of handling large volumes of data.
Develop reusable and maintainable Spark/PySpark frameworks and data-processing components.
Optimize Spark jobs for performance, scalability, memory utilization, and execution efficiency.
Work with complex transformations, joins, aggregations, partitioning, and large datasets.
Implement data validation, quality checks, error handling, and monitoring within data pipelines.
Work with AWS data services including EMR, Glue, S3, and Redshift.
Develop data pipelines supporting data lakes, warehouses, analytics, and downstream applications.
Troubleshoot production data pipeline and Spark processing issues.
Identify and resolve performance bottlenecks in Spark/PySpark workloads.
Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.
Must-Have Skills:
5+ years of hands-on experience in Data Engineering.
Strong hands-on experience with Apache Spark.
Strong hands-on experience with PySpark.
Strong programming experience in Python.
Proven experience developing and optimizing large-scale ETL/ELT pipelines.
Strong understanding of distributed computing and data-processing concepts.
Experience working with large datasets and complex data transformations.
Strong understanding of Spark performance optimization and tuning.
Experience with cloud-based data engineering, preferably AWS.
Experience with Amazon S3 and at least one AWS data-processing service such as EMR or Glue.
Good to Have:
AWS EMR
AWS Glue
Apache Airflow / MWAA
AWS Step Functions
Amazon Redshift
Hadoop ecosystem
Experience with data lake and data warehouse architectures.
Experience with CI/CD and production deployment of data pipelines.
AWS Certified Data Engineer or another relevant AWS certification.
Interview Process
Application review
5–10 minute initial screening call with the TA team
Technical interviews {Domain specific}
Practical test conducted in the presence of a panel member
Role match & offer