Job Summary
We are looking for a skilled Site Reliability Engineer (SRE) with strong expertise in Azure Monitoring and Observability. The candidate will be responsible for designing implementing and maintaining monitoring solutions to ensure high availability performance and reliability of cloud-based services on Microsoft Azure.
Responsibilities
Key Responsibilities
Design and implement endtoend monitoring solutions using Azurenative tools
Configure and manage
Azure Monitor
Log Analytics Workspaces
Application Insights
Azure Alerts & Action Groups
Develop dashboards and visualizations using
Azure Dashboards
Workbooks
Power BI optional
Define and refine
SLIs SLOs SLAs
Alerting strategies reduce alert fatigue
Perform proactive monitoring and incident management
Conduct
Root Cause Analysis RCA
Postincident reviews
Integrate monitoring with DevOps pipelines CICD
Automate monitoring configurations using
ARM templates Bicep Terraform
Work closely with development teams to ensure
Observability best practices are embedded in applications
Implement and manage
Logging frameworks
Distributed tracing