Job Title: Big Data Engineer
Experience: 5+Years
Employment Type: Full-Time
Work Mode: Remote
Job Summary
We are looking for an experienced Big Data Engineer with strong hands-on expertise in Scala, PySpark, AWS Glue, Big Data Pipeline Development, and Data Validation & Profiling.
The ideal candidate will be responsible for designing, developing, optimizing, and maintaining scalable data pipelines and data processing solutions on AWS. The role involves working with large and complex datasets, implementing data quality frameworks, supporting cloud data migrations, and optimizing data processing workloads.
The candidate should have strong programming and data engineering skills, along with experience working with modern big data technologies such as Databricks, Delta Lake, AWS EMR, Snowflake, and Parquet.
Key Responsibilities
- Design, develop, and maintain scalable Big Data pipelines using Scala, PySpark, and AWS Glue.
- Build robust and efficient ETL/ELT workflows for processing large volumes of structured and unstructured data.
- Develop distributed data processing solutions using Apache Spark and PySpark.
- Develop and maintain data ingestion, transformation, and integration pipelines across multiple data sources.
- Perform data validation, profiling, cleansing, and reconciliation to ensure data accuracy and completeness.
- Implement data quality checks and automated validation frameworks across data pipelines.
- Work with Databricks and Delta Lake for large-scale data processing and storage.
- Develop and optimize workloads running on AWS EMR.
- Work with Parquet and other distributed data formats to improve storage and processing efficiency.
- Support Snowflake migration initiatives, including data transformation, validation, and reconciliation.
- Identify and resolve data pipeline failures, data quality issues, and performance bottlenecks.
- Optimize Spark jobs, SQL queries, data partitions, joins, caching, and resource utilization.
- Implement monitoring, logging, error handling, and alerting for production data pipelines.
- Collaborate with data architects, analysts, application developers, QA teams, and business stakeholders.
- Participate in code reviews and follow coding, testing, documentation, and deployment best practices.
- Leverage AI coding assistants to improve development productivity, code quality, debugging, and documentation while following organizational security and governance standards.
- Contribute to continuous improvement of data engineering platforms, frameworks, and development processes.
Primary Skills
- Scala
- PySpark
- Apache Spark
- AWS Glue
- Big Data Pipeline Development
- ETL / ELT
- Data Validation
- Data Profiling
- Data Quality
Secondary Skills
- Databricks
- Delta Lake
- AWS EMR
- Snowflake Migration
- Parquet
- AI Coding Assistants
- Data Quality Frameworks
- Performance Optimization
- AWS Cloud
- SQL
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
- 4+ years of experience in Big Data Engineering or Data Engineering.
- Strong hands-on experience with Scala and/or PySpark.
- Strong experience developing and supporting large-scale data pipelines.
- Hands-on experience with AWS Glue and AWS data engineering services.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Experience with data validation, profiling, cleansing, and quality checks.
- Strong SQL skills and experience working with large datasets.
- Experience working with cloud-based data platforms and data lakes.
- Good understanding of data engineering concepts, ETL/ELT processes, partitioning, and data formats.
Preferred Qualifications
- Experience with Databricks and Delta Lake.
- Hands-on experience with AWS EMR.
- Experience with Snowflake and Snowflake migration projects.
- Strong understanding of Parquet and columnar storage formats.
- Experience developing or implementing data quality frameworks.
- Strong experience in Spark performance tuning and optimization.
- Experience with CI/CD and version control tools.
- Knowledge of AWS services such as S3, Lambda, CloudWatch, IAM, and Step Functions.
- Experience working in Agile/Scrum environments.
- Familiarity with modern AI coding assistants and AI-assisted software development.
Key Competencies
- Big Data Engineering
- Distributed Data Processing
- Data Pipeline Development
- AWS Cloud Data Engineering
- Spark & PySpark
- Scala Development
- Data Quality & Validation
- Data Profiling
- Data Migration
- Performance Optimization
- Problem Solving
- Production Support
Work Location: Remote