Job Description – Data EngineerPosition
Data Engineer – Azure / Databricks
Experience
4–8+ Years
Employment Type
Full-Time
Department
Data Engineering / Data & Analytics
Job Summary
We are looking for an experienced Data Engineer with strong expertise in Databricks, PySpark, Delta Lake, Azure Data Factory (ADF), Azure Cloud, Python, SQL, ADLS Gen2, Airflow, Logic Apps, CI/CD, MLflow, and Data Observability.
The ideal candidate will be responsible for designing, developing, and maintaining scalable data pipelines and cloud-based data platforms. The candidate should have hands-on experience with Azure data services, distributed data processing, data lake architecture, workflow orchestration, DevOps practices, and data quality/observability.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT workflows using Azure and Databricks.
- Develop data processing solutions using PySpark and Python.
- Build and optimize data pipelines using Databricks and Delta Lake.
- Implement reliable and scalable Delta Lake architectures for data ingestion, transformation, and storage.
- Develop and manage data integration pipelines using Azure Data Factory (ADF).
- Use Azure Logic Apps to implement workflow automation and system integrations.
- Develop and manage workflow orchestration using Apache Airflow.
- Work with Azure Data Lake Storage Gen2 (ADLS Gen2) for scalable data storage and data lake solutions.
- Write complex and optimized SQL queries for data transformation and analysis.
- Implement data ingestion from multiple structured and unstructured sources.
- Design data pipelines supporting batch and, where required, near-real-time processing.
- Implement CI/CD pipelines for automated deployment of data engineering solutions.
- Integrate development, testing, and deployment processes into DevOps workflows.
- Implement MLflow for experiment tracking, model lifecycle management, and integration with data/ML workflows.
- Establish and maintain data observability practices to monitor pipeline health, data quality, freshness, and reliability.
- Monitor pipeline performance and troubleshoot failures, data issues, and processing bottlenecks.
- Optimize Spark jobs, Databricks workloads, SQL queries, and data pipelines for performance and cost efficiency.
- Collaborate with Data Scientists, ML Engineers, BI Developers, Architects, and business stakeholders.
- Ensure data security, governance, quality, and compliance across Azure data environments.
- Create technical documentation for data pipelines, workflows, architectures, and operational processes.
Required Technical SkillsCategoryRequired SkillsData EngineeringData Pipelines, ETL/ELT, Data IntegrationDatabricksDatabricks, PySparkData LakeDelta Lake, ADLS Gen2AzureAzure Cloud, Azure Data Factory, Logic AppsProgrammingPython, SQLOrchestrationApache AirflowDevOpsCI/CD, Automated DeploymentML/MLOpsMLflowData QualityData Observability, Data MonitoringOptimizationSpark Performance Tuning, Pipeline OptimizationDatabricks & PySpark Responsibilities
- Develop scalable data transformation pipelines using PySpark.
- Implement batch data processing using Databricks.
- Build reusable notebooks, jobs, and data processing frameworks.
- Optimize Spark workloads through appropriate partitioning, caching, joins, and data-processing strategies.
- Implement Delta Lake features such as ACID transactions, schema management, and reliable data pipelines.
- Troubleshoot Databricks jobs and optimize compute utilization.
Azure Data Engineering
- Develop and manage Azure Data Factory pipelines for data ingestion and transformation.
- Implement data integration between Azure services and external data sources.
- Work with ADLS Gen2 to build scalable cloud-based data lake solutions.
- Use Azure Logic Apps for event-driven workflows and business process automation.
- Implement secure and reliable data movement across Azure environments.
- Monitor and troubleshoot Azure data pipelines and services.
CI/CD & DevOps
- Implement CI/CD processes for data engineering code, notebooks, pipelines, and configurations.
- Use Git-based version control and automated deployment practices.
- Support development, testing, staging, and production environments.
- Automate deployment and release processes for Azure and Databricks solutions.
- Follow DevOps best practices for data engineering projects.
Data Observability
- Implement monitoring for data pipelines, datasets, and workflows.
- Track pipeline failures, data freshness, completeness, and quality.
- Identify data anomalies and pipeline issues proactively.
- Establish appropriate alerts and operational dashboards.
- Work with development and data teams to resolve data quality and reliability issues.
MLflow Responsibilities
- Support integration of MLflow with data and machine learning workflows.
- Implement experiment tracking and model lifecycle processes where applicable.
- Collaborate with Data Scientists and ML Engineers to operationalize ML workflows.
- Support reproducibility and version management of ML-related assets.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- 4–8+ years of professional experience in Data Engineering.
- Strong hands-on experience with Databricks and PySpark.
- Strong knowledge of Delta Lake and modern data lake architectures.
- Hands-on experience with Azure Data Factory (ADF).
- Experience with Azure Logic Apps and Azure Cloud.
- Strong programming skills in Python and SQL.
- Experience working with ADLS Gen2.
- Hands-on experience with Apache Airflow or similar workflow orchestration tools.
- Experience implementing CI/CD for data engineering solutions.
- Knowledge of MLflow and MLOps concepts.
- Experience with Data Observability and data quality monitoring.
- Strong understanding of ETL/ELT, data integration, and data pipeline architecture.
Preferred Qualifications
- Experience designing enterprise-scale Azure Databricks architectures.
- Knowledge of Azure DevOps, GitHub Actions, or similar CI/CD platforms.
- Experience with real-time or streaming data processing.
- Knowledge of data governance, security, and access management.
- Experience with Infrastructure as Code tools such as Terraform.
- Understanding of medallion architecture (Bronze, Silver, Gold).
- Experience working in Agile/Scrum environments.
- Azure certifications such as Azure Data Engineer Associate (DP-203) are a plus.
- Experience working on ML/data science platforms and production ML pipelines is an advantage.
Key Competencies
- Strong data engineering and programming skills
- Excellent knowledge of Azure cloud data services
- Strong problem-solving and analytical abilities
- Ability to design scalable and reliable data pipelines
- Strong understanding of distributed data processing
- Performance and cost optimization mindset
- Good understanding of DevOps and CI/CD practices
- Strong communication and collaboration skills
- Ability to work effectively with Data Scientists, ML Engineers, and business teams
Work Location: Remote