Job Description – Senior Data Engineer (Databricks + PySpark + ADF + Airflow)
Job Title: Senior Data Engineer
Experience: 7+ Years
Employment Type: Full-Time
Role Summary
We are seeking an experienced Senior Data Engineer to drive cloud data modernization initiatives and build scalable, reliable, and AI-ready data platforms. The ideal candidate will have strong hands-on expertise in Databricks, PySpark, Delta Lake, Azure Data Factory (ADF), Azure Logic Apps, and Apache Airflow, with a proven ability to design and manage enterprise-scale data pipelines and orchestration frameworks.
The role requires deep expertise in data engineering, workflow automation, pipeline reliability, cloud-native architectures, and end-to-end data lifecycle management. The candidate will collaborate with data architects, data scientists, analysts, and engineering teams to deliver high-performance data solutions.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using:
- Databricks (PySpark)
- Azure Data Factory (ADF)
- Azure Logic Apps
- Apache Airflow
- Build and manage end-to-end orchestration frameworks integrating Airflow DAGs with ADF and Logic Apps.
- Develop advanced workflow orchestration patterns including:
- Event-driven workflows
- Micro-batch processing
- Hybrid scheduling models
- Dependency-based pipeline execution
- Design and implement high-performance ETL/ELT pipelines using:
- Databricks
- PySpark
- Delta Lake architecture
- Develop scalable data transformation frameworks for large-volume structured and semi-structured datasets.
- Implement pipeline dependency management, monitoring, error handling, and operational reliability practices.
- Integrate data pipelines with Azure services including:
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Blob Storage
- Azure event-driven services
- Build and maintain data workflow automation using Azure Logic Apps and Airflow orchestration.
- Implement data observability solutions including:
- Pipeline monitoring
- Logging
- Alerting
- Failure detection and recovery mechanisms
- Optimize data pipelines for:
- Performance
- Scalability
- Reliability
- Cost efficiency
- Implement CI/CD practices for data pipelines using:
- GitHub Actions
- Source control workflows
- Automated deployment processes
- Ensure data quality, governance, security, and compliance across enterprise data platforms.
- Collaborate with Data Scientists and AI/ML teams to prepare data platforms for analytics, ML, and GenAI workloads.
- Support data platform modernization and migration initiatives.
- Mentor junior data engineers and establish best practices for pipeline development, orchestration, and operational excellence.
- Participate in architecture discussions, design reviews, and technical solution planning.
Required QualificationsEducation
- Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, or related discipline.
Experience
- 7+ years of experience in Data Engineering.
- Strong experience designing and implementing enterprise-scale data platforms.
- Proven experience building cloud-native data solutions and automated workflows.
- Experience managing complex data pipelines in production environments.
Must-Have Technical SkillsDatabricks & PySpark
Strong hands-on experience with:
- Databricks
- PySpark
- Delta Lake architecture
- Distributed data processing
- Spark optimization techniques
- Large-scale ETL/ELT development
Azure Data Engineering
Highly proficient in:
- Azure Data Factory (ADF)
- Azure Logic Apps
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Blob Storage
- Azure cloud-native data architectures
Workflow Orchestration
Strong experience with:
- Apache Airflow
- DAG development
- Workflow scheduling
- Pipeline dependency management
- Job orchestration
- Failure handling and recovery strategies
Programming & Database Skills
Strong programming skills in:
Experience with:
- Data transformation frameworks
- Query optimization
- Data processing automation
ETL/ELT & Data Architecture
Strong understanding of:
- ETL/ELT design patterns
- Data pipeline architecture
- Data Lakehouse concepts
- Data modeling
- Data quality frameworks
- Data governance practices
Preferred Qualifications
- Experience with event-driven architecture and API-based integrations.
- Exposure to AI/ML data engineering workflows.
- Experience supporting ML pipelines and MLflow.
- Knowledge of Data Lakehouse architecture and governance frameworks.
- Experience with Azure Databricks ecosystem.
- Familiarity with monitoring and observability tools.
- Azure or Databricks certifications preferred.
- Experience working in Agile/Scrum environments.
Key Competencies
- Databricks Engineering
- PySpark Development
- Azure Data Engineering
- Data Pipeline Architecture
- Workflow Orchestration
- Airflow DAG Development
- ETL/ELT Development
- Delta Lake Architecture
- Cloud Data Modernization
- Pipeline Reliability Engineering
- Data Governance & Security
- Performance Optimization
- Technical Leadership
Role Outcomes
The successful candidate will:
- Build scalable and reliable cloud data platforms.
- Deliver high-performance Databricks and PySpark-based data solutions.
- Establish robust orchestration frameworks using Airflow, ADF, and Logic Apps.
- Enable AI/ML-ready data ecosystems through modern engineering practices.
- Improve operational reliability, automation, and efficiency of enterprise data workflows.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)