Lucknow, Uttar Pradesh
Job Summary
Job Responsibilities : Monitoring & Event Management Lead the Command Center monitoring team in a 24x7 operational environment. Monitor infrastructure, applications, cloud services, networks, and business-critical platforms. Ensure timely triage, analysis, and escalation of alerts and events. Validate monitoring effectiveness and reduce alert fatigue through continuous improvement. Ensure proper shift handovers and operational continuity. Major Incident Management Act as the Major Incident Lead during Priority 1 and Priority 2 incidents. Facilitate incident bridge calls and coordinate recovery efforts across support teams. Drive rapid restoration of services and minimize business impact. Manage escalations and ensure the engagement of appropriate technical teams. Maintain accurate incident timelines and documentation throughout the incident lifecycle. Communication & Stakeholder Management Provide timely and accurate updates to business users, support teams, and management. Communicate incident status, risks, actions, and recovery progress. Escalate unresolved issues as per defined escalation procedures. Support executive communications during critical service outages. Operational Excellence Ensure compliance with Incident, Major Incident, Event, and Problem Management processes. Conduct post-incident reviews and contribute to root cause analysis activities. Track and follow up on corrective and preventive actions. Identify recurring issues and recommend service improvement initiatives. Team Coordination Guide and mentor Command Center Analysts and Monitoring Specialists. Support training, knowledge sharing, and skill development within the team. Assist in resource planning and shift scheduling. Foster a culture of accountability, ownership, and operational excellence. Reporting & Metrics Generate daily, weekly, and monthly operational reports. Track and report on incident volumes, SLA performance, MTTR, and service availability. Maintain dashboards and operational metrics. Highlight trends, risks, and opportunities for improvement.
Key Responsibilities
Job Description : Key Competencies\\\\r\\\\nStrong communication and stakeholder management skills.\\\\r\\\\nAbility to handle high-pressure situations and critical incidents.\\\\r\\\\nExcellent coordination and decision-making skills.\\\\r\\\\nStrong analytical and problem-solving capability.\\\\r\\\\nCustomer-focused mindset.\\\\r\\\\nLeadership and team collaboration skills.\\\\r\\\\nKey Performance Indicators (KPIs)\\\\r\\\\nMean Time to Restore Service (MTTR)\\\\r\\\\nSLA Compliance\\\\r\\\\nService Availability/Uptime\\\\r\\\\nIncident Escalation Response Time\\\\r\\\\nMajor Incident Resolution Effectiveness\\\\r\\\\nMonitoring Alert Accuracy\\\\r\\\\nRepeat Incident Reduction\\\\r\\\\nStakeholder Satisfaction\\\\r\\\\nTimely Closure of Post-Incident Actions
Skill Requirements
Skill Requirement : echnical Skills Strong understanding of IT Infrastructure, Applications, Networks, Cloud, and Monitoring Operations. Hands-on experience with monitoring tools such as: Dynatrace Splunk Datadog SolarWinds AppDynamics SCOM Nagios Zabbix Knowledge of Azure, AWS, or other cloud platforms. Experience in incident coordination and service restoration activities.
Other Requirements
Other Requirement : Service Management Skills Major Incident Management Event Monitoring & Management Incident Management Problem Management Escalation Management Service Operations
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-