Project Role : IT Service Management Representative
Project Role Description : Manage delivery for IT production systems and services. Support the infrastructure service management lifecycle, ensuring client satisfaction and management of risk to services. Provide operational support, ensuring production systems and devices are online and available.
Must have skills : Critical Incident Management
Good to have skills : NA
Minimum 7.5 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
Lead the recovery of critical business services across a global enterprise environment. As the Incident Commander, you will drive rapid service restoration, lead cross-functional response teams, and leverage AI-driven insights to improve resilience, prevent disruptions, and enhance customer experience
Roles & Responsibilities:
Key Responsibilities
Incident Command & Service Recovery
Lead end-to-end management of P1/P2 incidents from declaration to closure.
Direct war rooms, executive communications, and crisis response activities.
Drive rapid restoration while balancing business impact, customer experience, and operational risk.
AI-Driven & Proactive Operations
Leverage Agentic AI, AIOps, and observability platforms for proactive risk detection and mitigation.
Drive automation and self-healing capabilities to reduce MTTR and improve service reliability.
Partner with Engineering, SRE, and Platform teams to strengthen operational resilience.
Customer & Stakeholder Leadership
Serve as the single point of leadership during major incidents.
Provide clear, business-focused communication to customers, executives, and stakeholders.
Foster strong collaboration across Infrastructure, Cloud, Security, Applications, and Vendor teams.
Continuous Improvement
Lead RCA, problem management, and corrective action governance.
Drive process optimization, operational excellence, and knowledge management.
Develop runbooks, playbooks, and AI-assisted restoration procedures.
Required Experience & Skills
10+ years of IT Service Management experience with significant Major Incident Management leadership.
Proven command-and-control leadership in high-pressure, business-critical environments.
Strong knowledge of ITIL, ServiceNow, Infrastructure, Cloud, Network, Security, and Applications.
Exposure to Agentic AI, AIOps, Automation, Observability, and SRE practices.
Strong communication, stakeholder management, decision-making, and problem-solving skills.
Education & Certifications
Education
Bachelor's degree in Information Technology, Computer Science, Engineering, or a related field.
Master's degree (preferred).
Preferred Certifications
ITIL Foundation / Managing Professional
ServiceNow ITSM Certification
Azure, AWS, or Google Cloud Certification
Agile / Scrum / SAFe Certification
SRE, AIOps, or AI-related certifications
Success Measures
Reduced MTTR and business-impacting outages.
95 compliance with incident response and communication SLAs.
Improved service availability, customer satisfaction, and operational resilience.
Reduced recurring major incidents through effective RCA and proactive prevention.
Increased adoption of AI-assisted, automated, and self-healing recovery capabilities
15 years full time education