Job Title
Big Data Engineer
Location
Noida, Uttar Pradesh
Work Mode
Remote
Employment Type
Full-time
Job Overview
We are seeking an experienced Big Data Engineer with strong expertise in Scala, PySpark, AWS Glue, Big Data Pipeline Development, and Data Validation & Profiling.
The ideal candidate will be responsible for designing, developing, and optimizing scalable big data pipelines and data processing solutions. You will work with modern cloud data technologies and contribute to data migration, data quality, validation, and performance optimization initiatives.
Key Responsibilities
- Design, develop, and maintain scalable big data pipelines using Scala, PySpark, and AWS Glue.
- Develop robust ETL/ELT workflows for large-scale data processing.
- Build and optimize distributed data processing solutions using Apache Spark.
- Perform comprehensive data validation, profiling, and quality checks across data pipelines.
- Develop reusable frameworks for data ingestion, transformation, validation, and processing.
- Work with Databricks, Delta Lake, AWS EMR, Snowflake, and other modern data platforms.
- Support Snowflake migration and cloud data modernization initiatives.
- Process and optimize data stored in formats such as Parquet.
- Implement data quality frameworks and automated validation processes.
- Identify and resolve data inconsistencies, pipeline failures, and performance bottlenecks.
- Optimize Spark jobs, data pipelines, queries, and resource utilization for performance and scalability.
- Collaborate with Data Architects, Data Engineers, Analysts, QA teams, and business stakeholders.
- Follow software engineering best practices including code reviews, version control, testing, and documentation.
- Leverage AI coding assistants to improve development productivity, code quality, and engineering efficiency.
- Monitor production pipelines and troubleshoot data processing issues.
Required Skills & Qualifications
- 4+ years of experience in Big Data Engineering or Data Engineering.
- Strong hands-on experience with Scala.
- Strong experience with PySpark and Apache Spark.
- Hands-on experience with AWS Glue.
- Strong experience in big data pipeline development.
- Experience with data validation and data profiling.
- Strong understanding of distributed data processing and ETL/ELT concepts.
- Experience working with large-scale datasets and cloud-based data platforms.
- Good understanding of data quality, data consistency, and validation techniques.
- Strong analytical and problem-solving skills.
Secondary / Preferred Skills
- Experience with Databricks.
- Hands-on experience with Delta Lake and Lakehouse architecture.
- Experience with AWS EMR.
- Experience supporting Snowflake migration projects.
- Strong knowledge of Parquet and other columnar data formats.
- Experience implementing data quality frameworks.
- Experience with automated data validation and reconciliation.
- Knowledge of performance optimization for Spark and big data workloads.
- Experience using AI coding assistants such as GitHub Copilot, Amazon Q Developer, ChatGPT, or similar tools.
- Familiarity with cloud-native data engineering practices.
- Experience with CI/CD and version control tools is a plus.
Key Responsibilities Around Data Quality
- Develop automated data validation and profiling processes.
- Validate source-to-target data mappings and transformations.
- Identify data anomalies, duplicates, missing values, and inconsistencies.
- Implement data quality rules and automated validation checks.
- Support reconciliation processes across data sources and targets.
- Monitor data quality metrics and continuously improve data reliability.
Performance Optimization
- Optimize PySpark and Scala Spark jobs for performance and scalability.
- Improve partitioning, caching, joins, and data processing strategies.
- Optimize AWS Glue and EMR workloads.
- Identify and resolve pipeline bottlenecks.
- Optimize storage and processing using Parquet and Delta Lake.
- Monitor resource utilization and improve overall pipeline efficiency.
Desired Candidate Profile
- Strong programming and analytical skills.
- Excellent understanding of distributed computing and big data technologies.
- Strong troubleshooting and debugging abilities.
- Good communication and collaboration skills.
- Ability to work independently in a remote environment.
- Comfortable working with large-scale data and complex engineering problems.
- Willingness to learn and adopt new technologies.
- Interest in using AI-assisted development tools to improve productivity.
Key Skills
Primary Skills:
Scala, PySpark, AWS Glue, Big Data Pipeline Development, Data Validation & Profiling
Secondary Skills:
Databricks, Delta Lake, AWS EMR, Snowflake Migration, Parquet, AI Coding Assistants, Data Quality Frameworks, Performance Optimization
Work Location: Hybrid remote in Remote