Site Reliability Engineer (SRE) – Data Reliability & Platform Engineering
Location: Bangalore (Whitefield)
Work Mode: Work from Office (5 Days)
Experience: 5+ Years
Job Summary
We are looking for a skilled Site Reliability Engineer (SRE) to join our Platform Engineering team. In this role, you will ensure the reliability, scalability, and performance of cloud platforms and data pipelines. You will work primarily with Google Cloud Platform (GCP), focusing on infrastructure automation, observability, CI/CD, and cloud-native operations while improving system uptime and data reliability.
Key Responsibilities
- Build and maintain highly available, scalable cloud infrastructure on GCP.
- Develop and automate infrastructure using Terraform and Infrastructure as Code (IaC).
- Design and manage CI/CD pipelines and GitOps workflows.
- Implement monitoring, logging, and observability using Prometheus, Grafana, Cloud Monitoring, and OpenTelemetry.
- Participate in incident response, root cause analysis (RCA), and production support.
- Ensure reliability and performance of batch and streaming data pipelines.
- Monitor data quality, freshness, lineage, and availability using modern data reliability practices.
- Collaborate with platform, application, and data engineering teams to improve cloud operations and automation.
- Optimize cloud infrastructure, deployment processes, and operational efficiency.
Required Skills
- 5–8+ years of experience in Site Reliability Engineering, Platform Engineering, or Data Engineering.
- Strong hands-on experience with Google Cloud Platform (GCP).
- Expertise in Terraform, CI/CD, GitOps, and Infrastructure as Code.
- Experience with Python, Go, Bash, or Shell scripting.
- Strong knowledge of Docker, Kubernetes, and distributed systems.
- Hands-on experience with Prometheus, Grafana, Cloud Monitoring, and OpenTelemetry.
- Good understanding of Linux, networking, and cloud infrastructure.
- Experience with BigQuery, Kafka, Spark, SQL/NoSQL databases, or modern data platforms.
- Knowledge of Ansible and cloud automation tools is a plus.
- Strong analytical, troubleshooting, and communication skills.
Preferred Skills
- Exposure to OCI or Azure.
- Experience with OpenLineage, data observability, and modern data reliability practices.
- Experience working in Agile/DevOps environments.
Application Question(s):
- Current CTC
- Expected CTC
- Currently serving notice period(Yes/ No) and mention the last working day?
Work Location: In person