Job Title
Senior Data Engineer – Databricks | PySpark | Azure Data Factory | Logic Apps | Airflow
Experience: 7+ Years
Employment Type: Full-Time
About the Role
We are seeking a highly skilled Senior Data Engineer to lead cloud data modernization initiatives and build scalable, AI-ready data platforms. The ideal candidate will have deep expertise in Databricks, PySpark, Azure Data Factory (ADF), Azure Logic Apps, and Apache Airflow, with a strong focus on workflow orchestration, pipeline reliability, and enterprise-scale data engineering.
You will be responsible for designing, developing, and optimizing robust data pipelines, implementing modern ETL/ELT architectures, and enabling reliable, high-performance data platforms that support analytics, AI/ML, and business intelligence initiatives.
Key Responsibilities
- Design, build, and maintain scalable data pipelines using Databricks (PySpark), Azure Data Factory (ADF), Azure Logic Apps, and Apache Airflow.
- Develop and manage end-to-end orchestration frameworks integrating Airflow DAGs, ADF pipelines, and Logic Apps.
- Implement advanced workflow orchestration patterns including event-driven, micro-batch, and hybrid scheduling.
- Ensure pipeline dependency management, execution reliability, fault tolerance, and operational excellence.
- Build high-performance ETL/ELT pipelines leveraging Databricks, PySpark, and Delta Lake architecture.
- Optimize data pipelines for performance, scalability, reliability, and cost efficiency.
- Implement monitoring, logging, alerting, and data observability across data workflows.
- Integrate pipelines with Azure services including ADLS Gen2, Azure Blob Storage, Event Grid, Event Hubs, and other Azure-native services.
- Develop reusable data engineering frameworks and automation for enterprise-scale data processing.
- Implement CI/CD pipelines for data engineering workflows using GitHub Actions, Azure DevOps, or equivalent tools.
- Ensure data quality through validation, governance, lineage, and security best practices.
- Collaborate closely with Data Science and AI teams to prepare high-quality datasets for Machine Learning and Generative AI workloads.
- Troubleshoot production issues, optimize pipeline performance, and ensure high system availability.
- Mentor junior engineers, conduct code reviews, and establish engineering best practices.
Required Skills & QualificationsExperience
- 7+ years of hands-on experience in Data Engineering.
- Proven experience designing enterprise-scale cloud data platforms.
Technical Skills
Must-Have
- Expert-level experience with Databricks (PySpark, Delta Lake).
- Highly proficient in Azure Data Factory (ADF).
- Highly proficient in Azure Logic Apps.
- Strong hands-on experience with Apache Airflow.
- Strong programming skills in Python.
- Advanced proficiency in SQL.
- Extensive experience building scalable ETL/ELT pipelines.
- Strong understanding of workflow orchestration and scheduling frameworks.
- Experience integrating multiple Azure data services.
- Strong knowledge of cloud-native data architectures on Microsoft Azure.
- Experience implementing monitoring, logging, and pipeline reliability solutions.
- Familiarity with Git-based version control and CI/CD automation.
Core Competencies
- Data Pipeline Development
- Workflow Orchestration
- Distributed Data Processing
- Performance Optimization
- Data Modeling
- Pipeline Automation
- Cloud Architecture
- Data Governance
- Problem Solving
- Cross-functional Collaboration
Preferred Qualifications
- Experience with event-driven architectures.
- Experience integrating REST APIs and enterprise applications.
- Exposure to AI/ML data pipelines and MLflow.
- Knowledge of Lakehouse Architecture and Delta Lake optimization.
- Experience with Data Governance and Data Catalog solutions.
- Azure Databricks Certification and/or Microsoft Azure Data Engineering Certification.
- Experience with Infrastructure as Code (Terraform/Bicep) is an added advantage.
Must-Have Skills
- Databricks
- PySpark
- Delta Lake
- Azure Data Factory (ADF)
- Azure Logic Apps
- Apache Airflow
- Python
- SQL
- ETL/ELT Pipeline Development
- Workflow Orchestration
- Azure Data Platform
Why Join Us?
- Work on enterprise-scale cloud data modernization initiatives.
- Build AI-ready, next-generation data platforms.
- Collaborate with experienced architects, data scientists, and engineering teams.
- Opportunity to work with cutting-edge Azure and Databricks technologies.
- Competitive compensation and excellent career growth opportunities.
Work Location: Hybrid remote in Noida, Uttar Pradesh