About the Role
We are looking for an experienced Application SRE with 4+ years of experience to ensure the high availability, scalability, and performance of our production applications. The ideal candidate will bridge the gap between application development and operations—leveraging expertise in application performance monitoring (APM), microservices architecture, CI/CD pipelines, cloud platforms, and incident management. This is a hybrid role with rotational shifts, including night shifts. Immediate joiners are preferred.
Skills RequiredApplication Site Reliability Engineering (App SRE) Microservices Architecture Application Performance Monitoring (APM) Root Cause Analysis (RCA) Prometheus / Grafana / Datadog / ELK CI/CD & Git AWS / Cloud Platforms Kubernetes & Docker Python / Bash Scripting Java / Node.js / Go (Basic Debugging) Incident Management Problem SolvingRequirements
- Education: Bachelor’s degree in Computer Science, IT, Software Engineering, or a related field.
- Experience: 4+ years of experience in Application SRE, Application Support Engineering, or Production Support.
- Application Knowledge: Strong understanding of application architectures (Microservices, REST APIs, Web Applications) and runtime environments.
- Observability & APM: Hands-on experience with Application Performance Monitoring (APM) tools (e.g., Datadog, Dynatrace, New Relic, AppDynamics) and logging stacks (ELK/EFK, Grafana Loki).
- Containers & Orchestration: Practical experience with Kubernetes and Docker for application deployment and troubleshooting.
- Cloud & Infrastructure: Hands-on experience with AWS or other major cloud platforms.
- CI/CD: Experience managing and supporting CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI) and Git workflows.
- Scripting & Automation: Strong scripting skills in Python, Bash, or Shell to automate application operational tasks.
- Problem Solving: Excellent debugging, code-level troubleshooting, and incident management skills.
- Work Flexibility: Willingness to work in a hybrid mode and rotational shifts, including night shifts.
- Availability: Immediate joiners preferred.
Nice To Have
- Experience with Caching & Messaging systems (Redis, Kafka, RabbitMQ) and Database performance tuning (PostgreSQL, MySQL, MongoDB).
- Solid knowledge of SLO, SLA, and SLI concepts applied at the service/API level.
- Familiarity with Infrastructure as Code (Terraform or Ansible).
- ITIL Foundation certification.
- AWS / Azure / GCP certifications.
Responsibilities
- Monitor, maintain, and optimize application health, performance, and availability across production environments.
- Troubleshoot and resolve complex L2/L3 application incidents, performance bottlenecks, and service disruptions.
- Perform deep-dive Root Cause Analysis (RCA) for application failures and drive permanent remediation.
- Track and improve application-level reliability metrics (SLOs, SLIs, error budgets, and MTTR).
- Manage incident, problem, and release management activities for software deployments.
- Support CI/CD pipelines, blue-green/canary deployments, and application rollout activities.
- Collaborate closely with Application Development, DevOps, and QA teams to fix bugs and enhance application resilience.
- Automate repetitive operational tasks (to reduce toil) and build automated self-healing scripts.
- Create and maintain operational documentation, runbooks, and application architecture diagrams.
Pay: ₹500,000.00 - ₹800,000.00 per year
Benefits:
Application Question(s):
- Night Shift along with rotational
- Immediate Joiners Only
Experience:
- Application support: 4 years (Required)
Work Location: In person