We are seeking a highly skilled and experienced Data Engineer with expertise in Databricks, Python, and AWS to join our dynamic team. The ideal candidate will be responsible for designing, developing, and maintaining scalable and robust data pipelines and data solutions within our cloud-native environment. This role requires a strong understanding of big data technologies, cloud architecture, and data warehousing principles to support various analytical and operational needs.
- Design, build, and optimize data pipelines using Databricks, Python, and AWS services to ingest, transform, and load large datasets from various sources.
- Develop and maintain ETL/ELT processes, ensuring data quality, reliability, and performance.
- Utilize Databricks platform features, including notebooks, Delta Lake, and Spark SQL, for data processing, analysis, and machine learning workflows.
- Implement and manage data solutions on AWS, leveraging services such as S3, EC2, Lambda, Glue, Redshift, Kinesis, and IAM.
- Write clean, efficient, and well-documented code in Python and PySpark for data manipulation, scripting, and automation.
- Collaborate with data scientists, analysts, and other engineering teams to understand data requirements and deliver appropriate data solutions.
- Monitor data pipelines and systems for performance, errors, and data integrity, implementing necessary alerts and troubleshooting procedures.
- Optimize existing data infrastructure and processes for cost efficiency, scalability, and improved performance.
- Ensure data security and compliance with industry best practices and company policies.
- Participate in code reviews, contribute to architectural discussions, and mentor junior team members.
- Stay up-to-date with emerging technologies and best practices in data engineering, cloud computing, and big data.