Job title: Azure DevOps - SRE Engineer
Shift: General Day shift
Experience: 4+ years in DevOps, SRE
Role Summary
We are looking for an Azure DevOps SRE Engineer to own the reliability, observability, and operational performance of production workloads on Microsoft Azure. This role blends platform engineering (CI/CD, Infrastructure as Code, Kubernetes) with Site Reliability Engineering practice — defining and defending SLOs, driving incident response and root-cause analysis, engineering away toil, and turning operational knowledge into automation and runbooks. You will work on actively operated production estates where reducing MTTA/MTTR and alert noise is the measure of success, including engagements where AI agents are being introduced into the incident lifecycle.
Key Responsibilities
- Reliability & SLO Ownership
- Incident Management & On-Call
- Observability Engineering
- Automation & Toil Reduction
- Platform & Delivery Engineering
Required Skills & Experience
● 4+ years in DevOps, SRE, or cloud platform engineering with production ownership of workloads on Microsoft Azure.
● Demonstrated SRE practice: has defined or operated against SLIs/SLOs, carried on-call responsibility, commanded or actively resolved high-severity incidents, and led postmortems.
● Hands-on observability engineering: metrics, logs, traces and alerting — building dashboards and alert rules that teams actually rely on, and cutting alert noise.
● CI/CD pipeline design and ownership using GitHub Actions and YAML workflows.
● Strong IaC skills with Terraform, including multi-environment promotion, reusable modules, and remote state management.
● Production Kubernetes operations (AKS preferred) and containerization with Docker — including scaling events, upgrades, and incident troubleshooting.
● Automation and scripting depth in Python, PowerShell, and/or Bash applied to operational tooling.
● Solid Azure platform fundamentals: networking (VNet/NSG), identity (Azure AD/Entra ID), and Key Vault-based secrets management.
● Systematic troubleshooting and debugging ability under production pressure, with structured root-cause analysis.
● Clear written and verbal communication — incident updates, stakeholder comms, and documentation that holds up in an audit or a handover.
Good to Have
Familiarity with the tooling in use across our current SRE engagements is an advantage. Equivalent experience on comparable platforms is acceptable — we care that you have operated a stack of this shape, not that you have used these exact products.
● ELK stack (Elasticsearch, Logstash, Kibana) for log aggregation and analysis — or equivalent (Splunk, Grafana Loki, Azure Log Analytics, Datadog).
● Azure Monitor and Application Insights for platform and application telemetry — or equivalent APM tooling (Dynatrace, New Relic, AppDynamics).
● Confluence for runbook and knowledge-base management — or equivalent documentation platform.
● JIRA for incident, RCA/RCCA, and corrective-action tracking — or equivalent (ServiceNow, Azure Boards, PagerDuty workflows).
● GitHub Actions as the delivery pipeline — or equivalent (Azure Pipelines, GitLab CI, Jenkins).
● Prometheus and Grafana for metrics and dashboarding.
● OpenTelemetry instrumentation and distributed tracing.
● Azure DevOps services (Boards, Repos, Pipelines, Artifacts).
● Exposure to AI/agentic operations tooling — LLM-assisted triage, automated RCA, or agent-driven remediation in a production or pilot setting.
● Databricks or data-platform reliability experience.
● Chaos engineering or fault-injection testing practice.
● FinOps and cloud cost anomaly management.
Certifications
● AZ-400: Designing and Implementing Microsoft DevOps Solutions — preferred for this role.
● AZ-104: Azure Administrator Associate.
● AZ-204: Azure Developer Associate — an acceptable alternative where the background is delivery-heavy.
● CKA (Certified Kubernetes Administrator) — good to have.
Pay: ₹700,000.00 - ₹900,000.00 per year
Work Location: In person