Project Role : Data Engineer
Project Role Description : Design, develop and maintain data solutions for data generation, collection, and processing. Create data pipelines, ensure data quality, and implement ETL (extract, transform and load) processes to migrate and deploy data across systems.
Must have skills : Databricks Unified Data Analytics Platform
Good to have skills : NA
Minimum
5 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
As a Data Engineer, a typical day involves designing, developing, and maintaining comprehensive data solutions that support the generation, collection, and processing of data. This role requires creating efficient data pipelines and ensuring the integrity and quality of data throughout its lifecycle. The position also involves implementing processes to extract, transform, and load data, facilitating seamless migration and deployment across various systems. Collaboration with different teams to optimize data workflows and support organizational data needs is a key aspect of daily activities, contributing to the overall data infrastructure and operational excellence.
Primary Technical Skills:
Databricks Notebooks - Experience in developing and managing Databricks notebooks (P1)
Python - For developing data processing workflows (P1)
Packages – Panda & Numpy (P1)
Apache Spark - Experience with Spark for big data processing, including Spark SQL and Spark Streaming.
ETL Tools - Knowledge of ETL processes and tools, specifically within the Databricks environment.
SQL - Strong skills for querying and managing data within Databricks SQL.
Performance Optimization- Skills in optimizing Spark jobs and troubleshooting performance issues.
Methodologies & Frameworks - Medallion architecture, Data Modelling (Star, Snowflake Schema), ETL, ELT, CI/CD, SDLC, Agile.
Secondary Technical Skills
ITIL incident, problem & change management
Visualization & Reporting - Power Bi
Cloud Platforms
- Familiarity with cloud platforms like AWS, Azure, or Google Cloud, since Databricks often integrates with cloud storage and services.
Soft Skills
Problem-Solving - Troubleshooting data issues and performance bottlenecks
Communication - Clear articulation of technical concepts to non-technical stakeholders
Collaboration:
Working effectively in cross-functional teams
Documentation - Maintaining clear and comprehensive documentation for pipelines, configurations, and processes
Roles & Responsibilities:
- Expected to be an SME, collaborate and manage the team to perform.
- Responsible for team decisions.
- Engage with multiple teams and contribute on key decisions.
- Provide solutions to problems for their immediate team and across multiple teams.
- Lead efforts to optimize data pipeline performance and scalability.
- Mentor junior team members to enhance their technical skills and project contributions.
- Coordinate cross-functional initiatives to align data engineering efforts with business objectives.
Professional & Technical Skills:
- Must To Have Skills: Proficiency in Databricks Unified Data Analytics Platform.
- Experience with building and managing scalable ETL pipelines.
- Strong knowledge of data integration, transformation, and migration techniques.
- Familiarity with cloud-based data storage and processing solutions.
- Ability to troubleshoot and resolve complex data-related issues efficiently.
- Understanding of data governance and quality assurance practices.
Additional Information:
- The candidate should have minimum 5 years of experience in Databricks Unified Data Analytics Platform.
- This position is based at our Bengaluru office.
- A 15 years full time education is required.