Job Title: Data Engineer
Experience: 5–8 Years
Employment Type: Full-Time
Position Summary
We are seeking an experienced Data Engineer to design, build, and optimize modern cloud-based data platforms that support analytics, AI/ML, and enterprise reporting. The ideal candidate will have hands-on expertise in Databricks, PySpark, Delta Lake, Azure Data Factory (ADF), Logic Apps, Apache Airflow, Python, SQL, ADLS Gen2, CI/CD, MLflow, Data Observability, and Azure Cloud. The role involves building scalable data pipelines, implementing modern data architectures, and ensuring reliable, high-quality data for business and AI applications.
Key Responsibilities
- Design, develop, and maintain scalable batch and streaming data pipelines using Databricks and PySpark.
- Build and optimize ETL/ELT workflows using Azure Data Factory (ADF), Logic Apps, and Apache Airflow.
- Develop robust data transformation pipelines using Python and SQL.
- Implement Delta Lake architecture for reliable, scalable, and high-performance data processing.
- Design and manage enterprise data lakes using Azure Data Lake Storage Gen2 (ADLS Gen2).
- Build cloud-native data solutions on Microsoft Azure.
- Optimize data processing jobs for performance, scalability, and cost efficiency.
- Develop and maintain CI/CD pipelines for automated deployment of data engineering solutions.
- Implement MLflow for machine learning experiment tracking, model lifecycle management, and integration with data pipelines.
- Monitor data quality, lineage, pipeline health, and reliability using Data Observability tools.
- Collaborate with Data Scientists, AI/ML Engineers, Business Analysts, and application teams to deliver trusted datasets.
- Implement security, governance, and compliance best practices across Azure data platforms.
- Troubleshoot production issues and continuously improve pipeline performance and reliability.
- Document architecture, data models, ETL processes, and operational procedures.
Required Skills
Technical Skills (Must Have)
- Databricks
- Apache Spark / PySpark
- Delta Lake
- Azure Data Factory (ADF)
- Azure Logic Apps
- Apache Airflow
- Python
- SQL
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Cloud
- ETL/ELT Development
- CI/CD
- Git
- MLflow
- Data Observability
Data Engineering Expertise
- Data Pipeline Development
- Batch & Streaming Data Processing
- Data Modeling
- Data Lake Architecture
- Data Warehousing
- Performance Optimization
- Data Quality Management
- Data Governance
- Workflow Automation
Good to Have
- Azure Synapse Analytics
- Azure Event Hubs
- Azure Functions
- Azure Key Vault
- Docker
- Kubernetes
- Terraform
- Kafka
- Power BI
- Snowflake
- Microsoft Fabric
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- 5–8 years of experience in Data Engineering and cloud-based data platforms.
- Strong hands-on experience with Databricks, PySpark, Delta Lake, and Azure Data Factory.
- Experience building scalable ETL/ELT pipelines on Azure Cloud.
- Proficiency in Python and SQL for large-scale data processing.
- Experience implementing CI/CD pipelines and Data Observability practices.
- Strong understanding of distributed computing, cloud architecture, and performance optimization.
- Excellent analytical, communication, and problem-solving skills.
Preferred Qualifications
- Experience supporting AI/ML and Generative AI data platforms.
- Knowledge of Microsoft Fabric, Azure Synapse Analytics, and modern Lakehouse architecture.
- Experience with MLOps practices using MLflow.
- Microsoft Azure Data Engineer Associate (DP-203) or Databricks certification is preferred.
- Experience working in Agile/Scrum environments.
What We Offer
- Opportunity to build enterprise-scale cloud data platforms and modern Lakehouse architectures.
- Exposure to the latest Azure, Databricks, AI, and Data Engineering technologies.
- Collaborative and innovation-driven work environment.
- Career growth through challenging cloud transformation and analytics initiatives.
- Competitive compensation and comprehensive employee benefits.
Work Location: Hybrid remote in Noida, Uttar Pradesh