Job Summary
We are looking for a skilled Data Engineer to design, develop, and maintain scalable data pipelines and cloud-based data solutions. The ideal candidate should have hands-on experience with Azure Data Services, Databricks, PySpark, Delta Lake, and modern data engineering practices. You will work closely with data architects, analysts, and business stakeholders to build reliable, high-performance data platforms that support analytics and business intelligence initiatives.
Key Responsibilities
- Design, develop, and optimize scalable ETL/ELT data pipelines using Azure Databricks and PySpark.
- Build and maintain data ingestion and transformation workflows using Azure Data Factory (ADF).
- Implement and manage Delta Lake architecture for reliable, scalable, and ACID-compliant data processing.
- Develop data orchestration workflows using Apache Airflow and Azure Logic Apps.
- Store, manage, and optimize data in Azure Data Lake Storage Gen2 (ADLS Gen2).
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop Python scripts for data processing, automation, and workflow orchestration.
- Build and manage CI/CD pipelines for data engineering solutions using Azure DevOps or GitHub Actions.
- Implement data quality checks, monitoring, and Data Observability practices to ensure data reliability.
- Integrate and manage MLflow for machine learning experiment tracking and model lifecycle management.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Collaborate with data scientists, BI developers, and cross-functional teams to deliver high-quality data solutions.
- Troubleshoot and resolve data pipeline failures, performance bottlenecks, and production issues.
- Ensure compliance with security, governance, and best practices within the Azure Cloud ecosystem.
- Maintain technical documentation for data pipelines, workflows, and architecture.
Required Skills
- Strong experience with Azure Databricks and PySpark.
- Hands-on experience with Delta Lake architecture and data engineering concepts.
- Expertise in Azure Data Factory (ADF) for ETL/ELT pipeline development.
- Experience with Azure Logic Apps and workflow automation.
- Strong knowledge of Apache Airflow for workflow orchestration.
- Proficiency in Python and SQL.
- Experience working with Azure Data Lake Storage Gen2 (ADLS Gen2).
- Knowledge of CI/CD implementation using Azure DevOps or GitHub Actions.
- Experience with MLflow for machine learning lifecycle management.
- Familiarity with Data Observability, monitoring, and data quality tools.
- Good understanding of Azure cloud services and distributed data processing.
- Strong analytical, debugging, and problem-solving skills.
- Excellent communication and collaboration abilities.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- Microsoft Azure certifications (preferred).
- Experience with data warehousing concepts and dimensional modeling.
- Familiarity with Power BI, Synapse Analytics, or Azure Fabric is an added advantage.
- Knowledge of Agile/Scrum development methodologies.
Preferred Technical Stack
- Cloud: Microsoft Azure
- Data Engineering: Azure Databricks, PySpark, Delta Lake
- Storage: Azure Data Lake Storage Gen2 (ADLS Gen2)
- ETL: Azure Data Factory (ADF)
- Workflow Orchestration: Apache Airflow, Azure Logic Apps
- Programming: Python, SQL
- Version Control & CI/CD: Git, Azure DevOps, GitHub Actions
- MLOps: MLflow
- Monitoring: Data Observability tools
- Methodology: Agile/Scrum
Work Location: Remote