As a Staff Engineer – Site Reliability Engineering, you will play a key technical leadership role in designing, building, and evolving our hybrid infrastructure and developer platforms. You’ll work hands-on across cloud automation, high-performance bare-metal systems, and DevOps tooling to deliver reliable, scalable infrastructure that accelerates software delivery. This role is highly cross-functional. You will collaborate closely with Development, DevOps, and QA teams to solve infrastructure and release challenges, support delivery cycles, and continuously improve platform capabilities. You’ll balance strong technical execution with architectural influence, contributing to system design while remaining directly involved in implementation and operational support. Occasional in-person meetings or team events may be required.
Ensure system availability and reliability through automated monitoring strategies
Produce post-mortems and implement resulting process improvements
Proactively mitigate operational risks through risk assessment and wider collaboration with
engineering teams
Design, implement, and continuously measure and improve risk mitigation strategies
Monitor system health through observability and telemetry
Unblock bottlenecks in system performance
Minimize emergency response time priods
Maintain internal tooling surrounding bug tracking, CI/CD pipelines, and wider
communication with the teams
Partner with Dev, DevOps, and QA teams to resolve infrastructure or deployment blockers during release cycles
Provide technical guidance and mentorship to platform engineers
Participate in architectural reviews, release readiness checkpoints, and root-cause analyses
10+ years of experience in infrastructure, platform, or DevOps engineering roles
Strong programming skills (e.g., Python, Go)
Hands-on experience with hybrid infrastructure (cloud + bare metal)
Deep knowledge of infrastructure-as-code tools (Terraform, Helm, Kubernetes)
Proven cross-functional collaboration skillsPreferred Qualifications
Experience with GPU-accelerated compute or HPC-style infrastructure
Familiarity with platform engineering or developer experience optimization
Experience with high-velocity release cycles and incident response
Experience with GPU-accelerated compute or HPC-style infrastructure
Familiarity with platform engineering or developer experience optimization
Experience with high-velocity release cycles and incident response