Excellent Opportunity Apply Now
We're looking for a Senior Site Reliability Engineer to take ownership of reliability, performance, and operational excellence across our cloud infrastructure and Kubernetes platforms. You'll lead the response to complex, high-severity incidents, raise the bar on automation and observability, and mentor other engineers as the team scales.
What You'll Do
- Lead troubleshooting and resolution of complex, high-severity incidents across AWS infrastructure and Amazon EKS clusters, often serving as incident commander for major incidents.
- Design and build Python-based automation frameworks and tooling that eliminate manual toil and improve reliability at scale, not just one-off scripts.
- Architect and maintain Terraform-based infrastructure-as-code, establishing reusable, secure, and scalable patterns across AWS environments.
- Drive observability strategy — design Grafana dashboards and alerting frameworks that surface the right signals to the right teams at the right time.
- Use SQL and CloudWatch Logs Insights to perform deep root-cause analysis on complex, cross-service incidents.
- Own end-to-end ITSM processes — Incident Management and Problem Management — including root cause analysis, post-incident reviews, and long-term remediation plans.
- Mentor junior and mid-level SREs, reviewing their troubleshooting approach, automation, and incident handling.
- Partner with engineering, product, and leadership teams to communicate incident impact, risk, and remediation plans clearly and confidently.
- Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.
- Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.
- Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.
- Participate in and provide senior-level escalation support for on-call rotations.
What We're Looking For
- 8 to 10 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, with a track record of owning reliability for production-critical systems.
- Deep, hands-on expertise with core and advanced AWS services (EC2, VPC, IAM, S3, RDS, CloudWatch, networking, etc.).
- Proven expertise troubleshooting complex Amazon EKS issues — cluster-level failures, networking, autoscaling, performance bottlenecks, and upgrade-related issues.
- Strong proficiency in Python for building automation frameworks, internal tooling, and operational systems.
- Extensive experience designing and maintaining Terraform modules and infrastructure patterns at scale.
- Strong command of ITSM frameworks, with hands-on ownership of Incident and Problem Management for high-severity issues.
- Advanced skills querying and analyzing data via SQL and AWS CloudWatch Logs Insights to drive root-cause analysis.
- Proven experience designing Grafana dashboards and alerting strategies that scale across multiple teams and services.
- Exceptional verbal and written communication skills — able to clearly articulate technical issues, risk, and remediation plans to engineering leadership and non-technical stakeholders alike.
- Experience mentoring or leading other engineers and contributing to team-level reliability strategy.
Mandatory Certifications
- AWS Certification — required (e.g., AWS Certified Solutions Architect – Professional, AWS Certified DevOps Engineer – Professional, or equivalent).
- Certified Kubernetes Administrator (CKA) or equivalent EKS/Kubernetes certification — required.
Soft Skills
- Calm, decisive leadership during high-pressure, high-severity incidents.
- A strong ownership mindset — drives issues to true resolution and follows through on long-term remediation.
- Natural mentor who raises the technical bar for the team.
- Collaborative cross-functional partner who works effectively with Dev, Infra, Product, and leadership.
Education
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent extensive practical experience.
Pay: ₹1,000,000.00 - ₹1,500,000.00 per year
Work Location: Hybrid remote in Bangalore City, Bengaluru, Karnataka