Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams.
Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide
Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise
Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up.
Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets
Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work
Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations.
This is a full-time role on a standard schedule, with participation in a global on-call rotation.