Mandatory Skills
- Scala/Python (PySpark)
- Big Data Knowledge
- Scripting (Shell/Python)
- GCP Data Tech Stack
- Apache Airflow
- CI/CD
- Apache Kafka
- Distributed Databases
- SQL/NoSQL
Key Responsibilities
- Design, develop, and maintain scalable data pipelines, frameworks, applications, and APIs using industry best practices.
- Build and monitor data systems and data quality processes to ensure reliability and performance.
- Conduct Proof of Concept (POC) and Proof of Technology (POT) evaluations.
- Review KPIs, monitoring metrics, and production alerts for critical systems.
- Collaborate with stakeholders to understand business requirements and deliver effective data solutions.
- Improve and automate data processes to enhance business value.
- Contribute to technical architecture discussions and maintain technical documentation.
- Stay updated with emerging technologies, tools, and methodologies in the data engineering space.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 8+ years of relevant experience in Scala/Python (PySpark), Kafka, and Distributed Databases.
- Strong understanding of Data Structures, Algorithms, SQL, and Query Optimization.
- Hands-on experience in building data ingestion and data consumption frameworks.
- Experience with orchestration tools such as Airflow.
- Strong experience in data processing and data manipulation.
- Expertise in GCP data processing tools and platforms, including GCS, Dataproc, BigQuery, Hive, and related services.
- Experience with stream processing using Kafka.
- Exposure to Lambda Architecture is a plus.
- Familiarity with reporting and visualization tools such as Tableau, Power BI, or Looker.
Pay: ₹120,000.00 - ₹150,000.00 per month
Benefits:
Work Location: Remote