Key Responsibilities
● Own production operations for the Saganova platform.
● Establish and maintain the Incident Management process.
● Define and govern P1/P2/P3/P4 severity and escalation criteria.
● Coordinate major production incidents and technical incident bridges.
● Ensure incidents are correctly triaged, assigned, escalated and closed.
● Establish incident communication and stakeholder update processes.
● Define and track operational SLAs/SLOs.
● Monitor platform availability, performance and production health.
● Work with Cloud/DevOps to define monitoring, alerts and dashboards.
● Ensure monitoring events are appropriately classified and actioned.
● Own operational Jira queues and incident tracking.
● Establish Problem Management for recurring production issues.
● Coordinate RCA and corrective actions following major incidents.
● Track RCA actions through to closure.
● Coordinate production readiness for releases.
● Maintain operational escalation matrices and on-call processes.
● Coordinate technology vendors during production incidents.
● Produce weekly/monthly operational performance reporting.
● Track availability, incidents, MTTA, MTTR, SLA breaches and recurring issues.
● Identify operational risks and improvement opportunities.
● Maintain production runbooks and operational documentation.
● Drive continuous improvement of production processes.
Required Skills
● Production Operations / Service Operations
● Incident Management
● Major Incident Management
● Problem Management
● SLA / SLO Management
● Production Monitoring & Alerting
● Operational Reporting
● Root Cause Analysis
● Escalation Management
● Jira / Service Management tools
● Cloud infrastructure fundamentals
● AWS fundamentals
● Application and infrastructure monitoring
● Logs and observability
● API and web application fundamentals
● Release / Change Management
● Stakeholder communication
● Vendor coordination
● Technical documentation
Pay: ₹60,000.00 - ₹80,000.00 per month
Work Location: In person