Role Summary
This role will lead a team responsible for monitoring, logging, tracing, and operational visibility across critical applications and infrastructure. The candidate will drive operational excellence, reliability, and continuous improvement while partnering closely with SREs, GCC and Security teams.
Key Responsibilities
-
Lead and mentor a team of observability engineers supporting enterprise platforms and services.
-
Define and execute the observability strategy, standards, and roadmap.
-
Oversee monitoring, logging, alerting, tracing, and dashboarding solutions.
-
Drive service reliability, incident response readiness, and operational excellence initiatives.
-
Collaborate with application, infrastructure, cloud, and SRE teams to improve system health and performance.
-
Establish KPIs, SLAs, and operational metrics to measure platform reliability and team effectiveness.
-
Manage hiring, performance development, resource planning, and stakeholder communications.
-
Ensure adoption of best practices for observability, automation, and proactive problem management.
-
Champion the adoption of AI and Copilot-enabled workflows within the Observability organization.
-
Evaluate and implement AI-driven monitoring, alert correlation, and incident management capabilities.
-
Define and execute the organization's strategy for AI, Agentic AI, and intelligent automation within the observability and operations ecosystem.
-
Drive the adoption of AI-powered operations, including autonomous incident management, intelligent alert correlation, predictive analytics, and self-healing platforms.
-
Lead initiatives leveraging Microsoft Copilot, Agentic AI frameworks, and AI agents to automate operational workflows, knowledge management, problem resolution, and service reliability improvements.
-
Partner with engineering teams to identify and prioritize use cases for Agentic AI that reduce manual effort and improve operational efficiency.
-
Partner with engineering and platform teams to build intelligent operational dashboards and automated remediation solutions.
Qualifications
-
Bachelor's degree in Computer Science, Engineering, or a related field.
-
8+ years of experience in infrastructure, operations, SRE, platform engineering, or observability domains.
-
3+ years of people management experience leading technical teams.
-
Strong knowledge of observability platforms such as Splunk, Datadog, AppDynamics, Dynatrace, Grafana, Prometheus, OpenTelemetry, or similar tools.
-
Experience working in cloud environments (Azure & GCP).
Experience leveraging Microsoft Copilot, Generative AI, and AI-powered observability capabilities to improve operational efficiency, incident response, and engineering productivity.
Knowledge of AI-assisted troubleshooting, anomaly detection, root cause analysis, and predictive monitoring solutions.
-
Excellent communication, stakeholder management, and leadership skills.
Preferred
-
Experience leading globally distributed teams.
-
Strong background in automation, DevOps, and reliability engineering practices.
-
Familiarity with enterprise-scale monitoring and incident management processes.
-
Experience with Generative AI, Agentic AI, Microsoft Copilot, Azure AI, OpenAI technologies, or similar AI platforms.