Job Description – Big Data Engineer
Job Title: Big Data Engineer
About the Role
We are looking for a skilled Big Data Engineer with strong experience in Scala, PySpark, AWS Glue, Databricks, Delta Lake, and Amazon EMR to design, develop, and optimize scalable data engineering solutions. The ideal candidate should have hands-on experience in big data pipeline development, data validation and profiling, data quality frameworks, performance optimization, and Snowflake migration.
The candidate will work on high-volume data processing, cloud-based data platforms, migration initiatives, and building reliable data pipelines using modern Big Data technologies.
Key Responsibilities
- Design, develop, and maintain scalable Big Data pipelines using Scala, PySpark, AWS Glue, and Databricks.
- Develop robust ETL/ELT pipelines for structured, semi-structured, and large-volume datasets.
- Work with Apache Spark/PySpark for distributed data processing and transformation.
- Develop and optimize data processing solutions using Delta Lake and Parquet.
- Build and manage data pipelines using AWS Glue and Amazon EMR.
- Perform data validation, profiling, reconciliation, and quality checks across data pipelines.
- Implement and maintain Data Quality Frameworks to ensure data accuracy, completeness, consistency, and reliability.
- Participate in Snowflake migration projects, including data extraction, transformation, loading, validation, and reconciliation.
- Optimize Spark jobs, SQL queries, data transformations, partitioning, caching, and resource utilization for improved performance.
- Troubleshoot production data pipelines and resolve data quality, performance, and processing issues.
- Implement appropriate error handling, logging, monitoring, and recovery mechanisms within data pipelines.
- Work with Delta Lake features such as ACID transactions, schema evolution, partitioning, and optimization.
- Utilize AI Coding Assistants to improve development productivity, code quality, debugging, documentation, and testing.
- Collaborate with data architects, analysts, application teams, and cloud engineers to deliver end-to-end data solutions.
- Follow best practices for coding, version control, CI/CD, security, and data governance.
Required Technical SkillsBig Data & Programming
- Strong hands-on experience with Scala and PySpark.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Experience developing enterprise-grade Big Data pipelines.
- Strong SQL skills for data transformation and validation.
AWS & Cloud
- Hands-on experience with AWS Glue.
- Experience with Amazon EMR and Spark-based workloads.
- Good understanding of AWS data services and cloud-based data engineering architectures.
Databricks & Delta Lake
- Strong experience with Databricks.
- Hands-on experience with Delta Lake, Delta tables, partitioning, schema evolution, and optimization.
- Understanding of Spark performance tuning within Databricks environments.
Data Engineering & Formats
- Experience with Parquet and other distributed data storage formats.
- Experience in ETL/ELT pipeline development.
- Strong understanding of data warehousing and data lake concepts.
- Experience with data validation and data profiling.
Data Quality
- Experience implementing Data Quality Frameworks.
- Knowledge of data completeness, accuracy, consistency, uniqueness, and reconciliation checks.
- Ability to identify and troubleshoot data anomalies.
Snowflake Migration
- Hands-on experience or strong understanding of Snowflake migration projects.
- Experience with data migration, transformation, validation, reconciliation, and performance optimization.
Performance Optimization
- Experience optimizing Spark jobs and Big Data pipelines.
- Knowledge of partitioning, caching, joins, shuffle optimization, file-size optimization, and query tuning.
- Ability to troubleshoot performance bottlenecks in distributed processing environments.
AI Coding Assistants
- Experience using AI-assisted coding tools such as GitHub Copilot or similar AI coding assistants.
- Ability to leverage AI tools for code generation, refactoring, debugging, test creation, and technical documentation.
Good to Have
- Experience with AWS S3, Lambda, Step Functions, CloudWatch, or related AWS services.
- Knowledge of Apache Airflow or other workflow orchestration tools.
- Experience with CI/CD tools such as Git, Jenkins, GitHub Actions, or Azure DevOps.
- Knowledge of data governance, metadata management, and data security.
- Experience working in Agile/Scrum environments.
- Exposure to Terraform/IaC and cloud automation.
- Experience with production support and monitoring of large-scale data platforms.
Key Competencies
- Strong problem-solving and analytical skills.
- Excellent understanding of distributed data processing.
- Ability to work with large and complex datasets.
- Strong debugging and performance-tuning capabilities.
- Good communication and collaboration skills.
- Ability to work independently and take ownership of data engineering deliverables.
- Strong focus on data quality, reliability, scalability, and performance.
Work Location: Remote