Job Description: DataBricks - Lead Programmer Analyst
We are looking for a highly self-motivated individual with DataBricks development as a Lead Programmer Analyst:
Required Skills & Qualifications:
-
Experience should have 5 to 7 Years of Data Engineering.
-
Expert-level Apache Spark skills using PySpark (Scala a plus).
-
Strong proficiency in SQL for data transformation and performance tuning.
-
Solid experience with Delta Lake and the Lakehouse/medallion architecture.
-
Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP) and its data services.
-
Experience with workflow orchestration (Databricks Workflows, Delta Live Tables, or Airflow).
-
Strong understanding of data modeling, data warehousing, and ETL/ELT design patterns.
-
Experience with Python for data engineering and automation.
-
Familiarity with CI/CD, Git, and DevOps practices for data.
-
Good understanding of data governance, security, and Unity Catalog.
-
Strong problem-solving and communication skills.
-
Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
Good To Have:
-
Databricks Certified Data Engineer Associate or Professional certification.
-
Experience with streaming technologies (Structured Streaming, Kafka, Event Hubs, Pub/Sub).
-
Exposure to machine learning workflows and MLflow.
-
Experience with Infrastructure-as-Code (Terraform) and cloud cost optimization.
-
Experience working in Agile delivery environments.
-
Experience in data warehouse design and maintenance
-
Experience in agile development processes using Jira and Confluence.
-
Experience in cross-functional teams.
Key Responsibilities:
-
Design, develop, and maintain scalable ETL/ELT data pipelines on the Databricks Lakehouse Platform.
-
Build and optimize Apache Spark (PySpark/Scala) jobs for batch and streaming data processing.
-
Implement and manage Delta Lake tables, including partitioning, schema evolution, and performance tuning (Z-ordering, caching, file compaction).
-
Develop and orchestrate workflows using Databricks Workflows, Delta Live Tables (DLT), and job scheduling tools such as Airflow.
-
Implement the medallion architecture (bronze, silver, gold layers) for reliable, governed data.
-
Integrate Databricks with cloud data services (AWS, Azure, or GCP) and source systems.
-
Apply data governance and security using Unity Catalog, access controls, and lineage.
-
Optimize cluster configuration, cost, and performance across workloads.
-
Collaborate with data analysts, data scientists, and business stakeholders to deliver trusted datasets.
-
Ensure data quality, testing, monitoring, and observability across pipelines.
-
Contribute to CI/CD practices for data engineering (version control, automated deployment).