Job Summary
We are seeking a highly skilled Data Engineer with expertise in Azure Cloud and modern data engineering technologies to build scalable, reliable, and high-performance data platforms. The ideal candidate should have hands-on experience with Azure Databricks, PySpark, Delta Lake, Azure Data Factory (ADF), Airflow, Logic Apps, Python, SQL, ADLS Gen2, MLflow, CI/CD, and Data Observability tools.
You will be responsible for designing, developing, and maintaining enterprise-grade data pipelines, enabling advanced analytics, machine learning, and business intelligence solutions across the organization.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Azure Databricks and PySpark.
- Build and optimize ETL/ELT workflows using Azure Data Factory (ADF).
- Implement and manage Delta Lake architecture to ensure reliable, scalable, and ACID-compliant data processing.
- Develop orchestration workflows using Apache Airflow and Azure Logic Apps.
- Manage structured and unstructured data in Azure Data Lake Storage Gen2 (ADLS Gen2).
- Develop efficient Python applications and optimized SQL queries for data processing and analytics.
- Create reusable data transformation frameworks and automate data workflows.
- Build and maintain CI/CD pipelines using Azure DevOps or GitHub Actions to automate deployment and testing.
- Implement MLflow for experiment tracking, model versioning, and machine learning lifecycle management.
- Monitor data pipelines using Data Observability tools to ensure data quality, reliability, lineage, and performance.
- Optimize Spark jobs and Databricks clusters for cost, scalability, and performance.
- Collaborate with Data Scientists, BI Developers, Architects, and Business Analysts to deliver data-driven solutions.
- Troubleshoot production issues, optimize workloads, and ensure high availability of data services.
- Implement security, governance, and access control policies across Azure data platforms.
- Document technical designs, architecture, workflows, and operational procedures.
Required Skills
- Strong hands-on experience with Azure Databricks.
- Expertise in PySpark for large-scale data processing.
- Experience implementing Delta Lake architecture.
- Proficiency in Azure Data Factory (ADF).
- Experience with Apache Airflow and Azure Logic Apps.
- Strong programming skills in Python.
- Advanced knowledge of SQL and database optimization.
- Experience working with Azure Data Lake Storage Gen2 (ADLS Gen2).
- Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions.
- Hands-on experience with MLflow.
- Knowledge of Data Observability, monitoring, logging, and alerting tools.
- Strong understanding of distributed computing and Spark optimization.
- Experience working in Microsoft Azure Cloud environments.
- Familiarity with Git version control and Agile development methodologies.
- Excellent analytical, troubleshooting, and communication skills.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- Microsoft Azure certifications such as Azure Data Engineer Associate (preferred).
- Experience with Azure Synapse Analytics, Microsoft Fabric, or Power BI is an added advantage.
- Knowledge of data warehousing, dimensional modeling, and data governance frameworks.
- Exposure to machine learning workflows and MLOps best practices.
The ideal candidate should possess hands-on experience with Microsoft Azure Cloud and modern data engineering technologies. Strong expertise in Azure Databricks, PySpark, and Delta Lake is essential for building scalable data processing solutions. Candidates should have experience developing ETL/ELT pipelines using Azure Data Factory (ADF) and orchestrating workflows with Apache Airflow and Azure Logic Apps.
Proficiency in Python and SQL is required for developing, transforming, and optimizing large-scale data solutions. Experience working with Azure Data Lake Storage Gen2 (ADLS Gen2) for data storage and management is expected.
The candidate should be familiar with CI/CD practices using Azure DevOps or GitHub Actions to automate deployments and ensure code quality. Knowledge of MLflow for machine learning lifecycle management and Data Observability tools for monitoring data quality, pipeline health, and system performance is highly desirable.
Additionally, experience with Git version control, distributed data processing, Spark optimization, Azure security best practices, and Agile/Scrum methodologies will be considered an advantage.
Work Location: Remote