Key Responsibilities
- Design, develop, and deploy data ingestion and transformation pipelines using Apache Spark (Scala).
- Work within Google Cloud Platform (GCP) components including Dataproc, BigQuery & Cloud Storage.
- Write complex SQL queries for data analysis, validation, and transformation in BigQuery.
- Develop Unix shell scripts for automation, data movement, and pipeline orchestration.
- Optimize and troubleshoot data pipelines for performance, scalability, and reliability.
- Collaborate with data engineers, analysts, and business stakeholders to understand requirements and ensure data quality.
- Contribute to code reviews, documentation, and CI/CD integration of data workflows.
Required Skills & Qualifications
- 4–6 years of hands-on experience in data engineering or related roles.
- Proven experience developing Spark applications in Scala.
- Strong experience working in Google Cloud Platform (GCP) ecosystem — including Dataproc, BigQuery, Cloud Storage.
- Proficient in SQL (especially BigQuery SQL dialects).
- Strong experience with Unix/Linux scripting for data automation.
- Familiarity with version control (Git) and CI/CD processes.
- Excellent problem-solving and debugging skills.
- Strong communication and documentation abilities.
Nice-to-Have Skills
- Experience with Python for data processing or automation.
- Knowledge of data governance, data quality, or metadata management best practices.
Education
- Bachelor’s degree in Computer Science, Engineering , or a related technical discipline (or equivalent work experience)
Pay: ₹1,800,000.00 - ₹3,400,000.00 per year
Work Location: Hybrid remote in Bengaluru, Karnataka (Bengaluru, Bengaluru Urban District)