Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration
- Job requisition ID : 109992
- Location: Bengaluru
- Entity: Deloitte Touche Tohmatsu India LLP
Location: Bangalore
Our Enterprise Technology & Performance team helps organizations build, operate, and optimize resilient cloud-native platforms. We are looking for an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate should possess strong expertise in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, preferably with experience in Banking & Financial Services, specifically Cards & Payments and Mobile Applications. Enterprise technology has to do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth. Learn more about: Customer
We are seeking an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, major incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate will be responsible for ensuring secure, scalable, highly available, and cost-effective cloud operations while driving operational excellence, service reliability, and continuous improvement across mission-critical enterprise applications. This role requires strong expertise in ITIL-based service management, Site Reliability Engineering (SRE), AWS cloud technologies, production support, automation, and stakeholder management. Experience in the Banking & Financial Services domain, particularly Cards & Payments, Mobile Applications, and Cloud-Native Solutions, will be highly advantageous.
Lead enterprise-wide Incident, Problem, and Change Management activities aligned with ITIL best practices.
Own the end-to-end lifecycle of production incidents (P1–P4), ensuring timely identification, escalation, communication, resolution, and closure.
Drive service recovery, Root Cause Analysis (RCA), Post Incident Reviews (PIR), and Corrective & Preventive Actions (CAPA) to improve service reliability and reduce MTTR.
Lead 24x7 production support operations, managing L2/L3 application and infrastructure support across Linux-based environments.
Oversee production deployments, release management, infrastructure operations, and Site Reliability Engineering (SRE) initiatives.
Manage Disaster Recovery (DR), High Availability (HA) architecture, and Active-Active/Active-Passive failover strategies.
Drive automation, cloud infrastructure optimization, capacity planning, and operational excellence using AWS, Kubernetes, Docker, and CI/CD pipelines.
Monitor and optimize production environments using observability tools including Grafana, Kibana, ELK, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7.
Support REST API-based applications, perform production troubleshooting using SQL, and collaborate with engineering teams to ensure highly resilient and secure production environments.
Lead and mentor SRE and Production Support teams while ensuring SLA, SLO, KPI, and customer satisfaction targets are consistently achieved.
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 10–15+ years of overall IT experience with at least 8+ years specializing in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE).
- Strong experience in Banking & Financial Services, preferably supporting Cards & Payments, Mobile Applications, and Cloud-Native solutions.
- Proven expertise in Incident, Problem & Change Management (ITIL), Major Incident Management, Service Recovery, Root Cause Analysis (RCA), and Production Operations.
- Experience leading 24x7 production support teams and managing L2/L3 application and infrastructure support across Linux environments.
- Strong knowledge of AWS services including EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, CloudWatch, and secure cloud architecture.
- Experience designing highly available, scalable, secure, and disaster recovery-enabled cloud solutions.
- Hands-on experience with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, AWS CloudFormation, Ansible), and automation using Python or Bash.
- Strong understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices.
- Experience with monitoring and observability tools including Grafana, Kibana, ELK Stack, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7.
- Proven ability to drive SLA/SLO/KPI compliance, reduce MTTR, improve service reliability, and implement automation and continuous improvement initiatives.
- Strong leadership, stakeholder management, client communication, analytical, and problem-solving skills with the ability to coordinate cross-functional teams during critical incidents.
- ITIL Foundation Certification is preferred.
- AWS Certified Solutions Architect – Associate or Professional certification is highly preferred.