Job Description
We are seeking a motivated Associate Site Reliability Engineer (L1) to join our Infrastructure team as the first line of defense for our production environments. In this role, you will monitor platform health, triage incoming alerts, execute standard operating runbooks, and assist in maintaining high availability across our cloud services. This position is ideal for an early-career technologist or IT professional looking to build a career in modern DevOps, Cloud Operations, and Site Reliability Engineering.
Responsibilities
Monitor system health and performance dashboards continuously using tools like Grafana, Datadog, and Prometheus.
Acknowledge, triage, and categorize real-time production alerts from PagerDuty and monitoring platforms.
Execute standardized runbooks and SOPs to remediate routine service outages, clear resource bottlenecks, and restart failing processes.
Escalate unresolved or high-severity incidents to L2/L3 SREs and development teams with detailed contextual diagnostics.
Document incident timelines, maintain accurate ticket records in Jira/ServiceNow, and keep status pages updated during outages.
Identify out-of-date documentation and suggest improvements to operational runbooks and troubleshooting guides.
Assist senior SREs with basic infrastructure maintenance, deployments, and routine system checks.
Requirements and Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent hands-on technical training/bootcamp.
0–2 years of experience in IT support, helpdesk, Network Operations Center (NOC), or technical operations.
Foundational knowledge of Linux/Unix command-line navigation, file systems, and log analysis (tail, grep, awk, top).
Basic understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS status codes, and SSH.
Familiarity with container concepts and command-line interfaces like Docker or Kubernetes (kubectl).
Strong analytical problem-solving skills with a methodical approach to troubleshooting.
Clear written and verbal communication skills for incident reporting and team handoffs.
Ability to work in shifts or participate in an entry-level on-call rotation schedule.
Nice to Have
Hands-on exposure to writing basic automation scripts in Python or Bash.
Foundational cloud certification (e.g., AWS Certified Cloud Practitioner, Azure Fundamentals, or GCP Digital Leader).
Basic understanding of Version Control Systems like Git and CI/CD deployment pipelines.
Prior experience using incident management tools like PagerDuty, Opsgenie, Jira, or ServiceNow.
Required Skills
Linux/Unix Administration, Monitoring & Observability (Grafana/Datadog), Alert Triage, Incident Response, Networking Fundamentals (TCP/IP, DNS, HTTP), Basic Docker/Containers, Incident Documentation.
Pay: ₹641,284.09 - ₹796,768.29 per year
Benefits:
Application Question(s):
Experience:
Work Location: In person