Job Title: Azure Data Engineer
Experience: 4–8 Years
Employment Type: Full-Time
Job Summary
We are looking for a highly skilled Azure Data Engineer with strong expertise in Databricks, PySpark, Delta Lake, Azure Data Factory (ADF), Logic Apps, and Apache Airflow. The ideal candidate will design, develop, and optimize scalable data pipelines, implement modern data lakehouse architectures, and enable reliable data integration across enterprise platforms.
Key ResponsibilitiesData Engineering & ETL Development
- Design, develop, and maintain scalable ETL/ELT pipelines using Azure Databricks and PySpark.
- Build and optimize data processing workflows for structured, semi-structured, and unstructured data.
- Implement data ingestion, transformation, and loading processes for enterprise data platforms.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
Data Lakehouse & Delta Lake
- Develop and maintain Delta Lake architecture using Bronze, Silver, and Gold layers.
- Implement Delta Lake features such as ACID transactions, schema evolution, time travel, and data versioning.
- Ensure high-quality, reliable, and consistent datasets.
Azure Data Factory (ADF)
- Design and manage Azure Data Factory pipelines for data movement and orchestration.
- Integrate data from multiple on-premises and cloud sources.
- Implement monitoring, logging, and error-handling mechanisms for ADF pipelines.
Workflow Automation
- Build workflow orchestration using Apache Airflow.
- Automate data processing workflows, scheduling, and dependency management.
- Develop automation using Azure Logic Apps for event-driven integrations and business workflows.
Cloud & Data Integration
- Integrate Azure services including Azure Storage, Azure SQL Database, Azure Synapse Analytics, Event Hub, and Azure Key Vault.
- Develop secure, scalable, and high-performance cloud data solutions.
- Collaborate with Data Scientists, BI Developers, and Software Engineers to support analytics and reporting.
Performance & Governance
- Optimize Spark clusters and Databricks workloads.
- Implement data quality checks, validation, and governance best practices.
- Monitor production pipelines and troubleshoot performance issues.
- Maintain documentation for architecture, pipelines, and deployment processes.
Required Skills
- Strong experience with Azure Databricks
- Expertise in PySpark
- Hands-on experience with Delta Lake
- Experience with Azure Data Factory (ADF)
- Strong knowledge of Apache Airflow
- Experience with Azure Logic Apps
- Good understanding of ETL/ELT concepts and data pipeline development
- Experience working with Azure Data Lake Storage (ADLS Gen2)
- Strong SQL skills
- Knowledge of Spark optimization and performance tuning
- Experience with Git-based version control
- Understanding of CI/CD for data engineering solutions
Preferred Skills
- Azure Synapse Analytics
- Azure SQL Database
- Azure Key Vault
- Azure Event Hub
- Azure DevOps CI/CD
- Python programming
- Power BI integration
- Data Modeling
- REST APIs
- Docker and Kubernetes (preferred)
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
- 4–8 years of experience in Data Engineering or Azure Data Platform development.
- Microsoft Azure certifications are a plus.
Nice to Have
- Experience with Medallion Architecture
- Knowledge of Data Governance and Data Quality frameworks
- Experience with Streaming data using Spark Structured Streaming or Kafka
- Familiarity with Snowflake or Microsoft Fabric is an added advantage
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)