12–18 Years
Bangalore / Hybrid
We are looking for an experienced Monitoring & Observability Delivery Manager to lead the delivery, operations, and continuous improvement of enterprise observability platforms across cloud and on-premises environments. The ideal candidate should possess strong expertise in logs, metrics, traces, dashboards, alerting, incident management, and SRE practices, along with proven experience in managing large delivery teams and customer engagements.
This role requires a combination of technical expertise, delivery leadership, stakeholder management, and operational excellence to ensure high platform availability, proactive monitoring, and rapid incident resolution.
Lead end-to-end delivery of Monitoring & Observability services across multiple customer environments.
Manage large teams of engineers, technical leads, and architects.
Drive project execution, planning, governance, and customer communication.
Ensure SLA, KPI, and operational excellence targets are consistently achieved.
Manage production support, incident management, problem management, and change management processes.
Present service health, operational metrics, and executive dashboards to customer leadership.
Design and implement enterprise monitoring strategies.
Build observability solutions using logs, metrics, traces, and distributed tracing.
Define monitoring standards, alerting strategies, dashboards, and operational KPIs.
Drive observability maturity across applications, infrastructure, cloud, and Kubernetes platforms.
Reduce alert noise through alert tuning and automation.
Strong understanding of:
Application Logs
Infrastructure Logs
Kubernetes Logs
Linux & Windows Logs
Database Logs
Middleware Logs
API Gateway Logs
Audit Logs
Security Logs
Ability to
Analyze log patterns
Troubleshoot production incidents
Identify recurring failures
Perform Root Cause Analysis (RCA)
Build centralized log analytics solutions
Experience in monitoring
CPU
Memory
Disk
Network
JVM
Containers
Kubernetes
Database Performance
API Performance
Application Availability
Response Time
Throughput
Error Rates
Capacity Planning
Infrastructure Health
Should be capable of defining
Golden Signals
RED Metrics
USE Metrics
SLI/SLO/SLA
Business KPIs
Hands-on knowledge of
Hands-on experience with one or more of the following
Dynatrace
Datadog
Splunk
Elastic (ELK)
Grafana
Prometheus
New Relic
AppDynamics
OpenTelemetry
Jaeger
Zipkin
Lead P1/P2 production incidents.
Drive Root Cause Analysis (RCA).
Identify recurring issues and preventive actions.
Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
Implement proactive monitoring and automation.
Experience with
Python
Shell Scripting
PowerShell
Terraform
Ansible
Jenkins
GitHub Actions
Automation of
Monitoring deployment
Alert creation
Dashboard provisioning
Health checks
Incident remediation
Work closely with customer leadership, architects, infrastructure teams, and application teams.
Present weekly/monthly operational reviews.
Manage escalations and executive communications.
Define roadmap for observability transformation.
12+ years of IT experience.
5+ years managing Monitoring & Observability platforms.
Strong understanding of logs, metrics, traces, and distributed systems.
Experience managing cloud-native environments and Kubernetes.
Strong troubleshooting and production support expertise.
Excellent communication, stakeholder management, and leadership skills.
Experience managing teams of 20+ engineers.
AWS/Azure/GCP Certification
Dynatrace Professional Certification
Datadog Certification
Splunk Certification
Kubernetes (CKA/CKAD)
ITIL Foundation
SRE or DevOps certifications
The ideal candidate is a technically strong delivery leader who can interpret logs, metrics, and traces to quickly identify production issues, lead cross-functional incident response, and build scalable observability solutions. They should have experience delivering enterprise monitoring services, mentoring technical teams, engaging with customers, and driving continuous improvements through automation and observability best practices.