Senior Site Reliability Engineer (SRE)
Location: Delhi NCR | Employment Type: Full-Time | Experience: 4+ Years
About the Role
We are looking for a proactive and detail-oriented Sr. Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. You'll monitor production systems, troubleshoot issues, and collaborate with cross-functional teams to drive faster resolution and continuous improvement — playing a key role in maintaining system stability and enhancing observability across our microservices-based platform.
Key Responsibilities
- Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
- Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
- Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
- Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
- Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
- Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
- Participate in incident/severity calls, ensuring clear communication and coordination
- Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting
Required Skills & Experience
- Strong understanding of Linux/Unix systems for application support
- Hands-on experience troubleshooting applications in staging and production environments
- Ability to monitor system performance and identify root causes using logs and metrics
- Experience working with Kubernetes and microservices-based architectures
- Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
- Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
- Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
- Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
- Experience with ticketing and documentation tools such as Jira and Confluence
- Minimum 4+ years of experience in application support or reliability engineering
- Bachelor's degree in Computer Science, Information Technology, or a related field
- Relevant certifications (Cloud, Kubernetes, Microservices) are a plus
Work Schedule
Willingness to work in a 24x7 environment, including weekends and on-call rotations.
Application Question(s):
- How many years of experience do you have in application support, DevOps, or site reliability engineering?
- Are you willing to work in a 24x7 environment, including weekends and on-call rotations?
- Location Delhi NCR Rotational are you comfortable ?
Experience:
- Linux/Unix system administration and troubleshooting: 4 years (Required)
Work Location: In person