We are seeking a motivated and detail-oriented Data Engineer with 2-3 years of experience in data engineering, ETL development, and cloud-based data solutions. The ideal candidate should have strong hands-on experience with SQL, Python, PySpark, Airflow, GitHub, and AWS services including Redshift, Glue, S3, Lambda, and EMR.
- Design, develop, and maintain scalable ETL/ELT pipelines using Python and PySpark.
- Build and optimize data workflows for processing large datasets.
- Develop and support cloud-based data solutions using AWS Redshift, Glue, S3, Lambda, and EMR.
- Write complex SQL queries and optimize data processing performance.
- Create and maintain workflows using Apache Airflow.
- Use GitHub for source code management and collaborative development.
- Investigate and resolve production issues related to data pipelines.
- Participate in code reviews and maintain technical documentation.
- Collaborate with cross-functional teams to deliver reliable data solutions.
- Identify opportunities for process improvement and automation.
- Strong working knowledge of SQL, Python, PySpark, Airflow, and GitHub.
- Experience with AWS services: Redshift, Glue, S3, Lambda, and EMR.
- Understanding of ETL/ELT concepts and data warehousing principles.
- Knowledge of data modeling and data quality practices.
- Strong analytical and problem-solving skills.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- 2 -3 years of experience in Data Engineering, ETL Development, or a related role.
- Experience working in Agile development environments is preferred.
- Exposure to CI/CD concepts and DevOps practices.
- Basic understanding of Docker and containerization.
- Familiarity with Databricks or modern data platforms.
- Knowledge of data governance and data security best practices