Site Reliability Engineer Pune · Full Time · Remote
About the Role The Site Reliability Engineer will be responsible for ensuring the reliability and performance of .NET applications and AWS infrastructure. The role involves partnering with application engineers, embedding reliability into new feature design, and optimizing infrastructure for consistent and reproducible environments. Success in this role means driving data-driven reliability improvements, reducing toil, and freeing the team for higher-impact engineering work.
Key Responsibilities
- Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.
- Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.
- Partner with application engineers to embed reliability into new feature design and deployment practices.
- Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.
- Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads.
- Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.
- Automate operational tasks, deployment pipelines, and disaster recovery procedures.
- Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.
- Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.
- Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.
- Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.
- Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly.
- Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.
- Participate in on-call rotation to respond to production incidents and drive swift resolution.
Requirements
- Have 4+ years of experience in SRE, DevOps, platform engineering, or a systems-focused software engineering role.
- Possess C#/.NET engineering ability — can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure.
- Have AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.
- Have RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.
- Have native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.
- Have AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.
- Have SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.
- Have strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.
- Have excellent communication skills and a collaborative, blameless engineering mindset.
Good to Have
- Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.
What We Offer
- Opportunity to work with a collaborative and blameless engineering team.
- Chance to drive data-driven reliability improvements and reduce toil.
- Freedom to experiment with new technology and tools to work smarter and faster.