The Autonomous Recovery Service (RCV) team is responsible for delivering highly available, secure, and resilient cloud services that protect Oracle Cloud Infrastructure (OCI) customer data. Our mission is to ensure customers can confidently recover from planned and unplanned events through industry-leading recovery capabilities, intelligent automation, and operational excellence.
As a Software Development & Site Reliability Engineer , you will play a critical role in operating and continuously improving mission-critical cloud services. This is an operations-first engineering role where you'll own the health, reliability, and lifecycle of production services while developing software and automation that reduce operational complexity, improve service resilience, and enhance the customer experience.
You'll work across the full service lifecycle—from deployment, patching, upgrades, monitoring, incident response, and root cause analysis to designing automation, improving observability, and implementing engineering solutions that eliminate repetitive operational work. Success in this role requires curiosity, strong analytical thinking, and a passion for solving complex operational challenges through software and automation.
Our engineers embrace AI as a force multiplier, using AI-assisted development and operational tools to accelerate problem solving, improve productivity, and build smarter, more autonomous systems. We value engineers who combine technical expertise with critical thinking, sound engineering judgment, and a continuous improvement mindset to challenge existing processes and drive innovation.
If you enjoy owning production services, building reliable cloud infrastructure, automating everything possible, and working on technology that protects mission-critical customer data at cloud scale, we'd love to have you join our team.