Project Role : Custom Software Engineer
Project Role Description : Develop custom software solutions to design, code, and enhance components across systems or applications. Use modern frameworks and agile practices to deliver scalable, high-performing solutions tailored to specific business needs.
Must have skills : Databricks Unified Data Analytics Platform
Good to have skills : NA
Minimum 5 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
As a Databricks Data Engineer, you will design, develop, and maintain scalable data solutions using the Databricks Unified Data Analytics Platform. The role requires expertise in data engineering, Apache Spark, PySpark, SQL, Delta Lake, and cloud-based data platforms to build reliable and high-performance data pipelines.
The engineer will develop ETL/ELT workflows, perform large-scale data transformations, optimize Spark workloads, and support enterprise analytics initiatives using modern cloud data technologies. The role involves collaborating with data architects, analysts, engineers, and business stakeholders to deliver secure, scalable, and efficient data solutions.
Roles & Responsibilities:
1. Design, develop, and maintain scalable data pipelines using Databricks and Apache Spark.
2. Develop ETL/ELT workflows using PySpark, Spark SQL, and SQL for enterprise data processing.
3. Build, implement, and optimize Delta Lake solutions for scalable and reliable data management.
4. Perform data ingestion, transformation, validation, and integration from multiple data sources.
5. Develop and manage Databricks notebooks, jobs, workflows, clusters, and production data processes.
6. Optimize Spark performance, cluster utilization, data processing efficiency, and pipeline execution.
7. Implement data quality checks, validation frameworks, monitoring, and troubleshooting processes.
8. Develop reusable data engineering components and follow best practices for scalable pipeline development.
9. Collaborate with data architects, data analysts, BI teams, and business stakeholders to deliver analytics solutions.
10. Support cloud-based data platforms and enterprise data modernization initiatives.
11. Implement CI/CD practices, version control, and DevOps standards for data engineering workflows.
12. Participate in Agile ceremonies including sprint planning, reviews, and continuous improvement activities.
Professional & Technical Skills:
1. Must Have Skills: Strong proficiency in Databricks Unified Data Analytics Platform.
2. Strong experience with Apache Spark, PySpark, Spark SQL, and distributed data processing.
3. Hands-on experience with Delta Lake, Lakehouse architecture, and modern data engineering patterns.
4. Strong SQL skills with experience in complex queries, optimization, and data transformation.
5. Experience designing and developing ETL/ELT data pipelines.
6. Experience working with cloud platforms including Microsoft Azure, AWS, or Google Cloud Platform (GCP).
7. Strong understanding of data warehousing concepts, dimensional modeling, and enterprise data architecture.
8. Experience with data ingestion frameworks and orchestration tools such as Apache Airflow, Azure Data Factory, or similar technologies.
9. Knowledge of data formats including Parquet, JSON, CSV, and other structured/semi-structured data formats.
10. Experience with data quality, validation, monitoring, and troubleshooting processes.
11. Understanding of Git, CI/CD pipelines, DevOps practices, and Agile methodologies.
12. Familiarity with cloud storage and data services such as Azure Data Lake, Amazon S3, Google Cloud Storage, or equivalent platforms.
13. Strong analytical, troubleshooting, communication, and problem-solving skills.
Additional Information:
1. The candidate should have minimum 6 years of experience in Databricks Unified Data Analytics Platform.
2. This position is based at our Hyderabad office.
3. 15 years full time education is required.
15 years full time education