### JD - Principal Platform Engineer (AIOps Specialist)
#### About the Role
We are looking for a highly seasoned and pragmatic **Principal Platform Engineer** to lead our platform & infrastructure strategy for Rakuten's Entertainment products and services.
With over 15 years of experience, you are a "full-stack" engineer who possesses deep expertise in both application development and infrastructure architecture. You operate with a high degree of autonomy, capable of navigating complex codebases and cloud architectures simultaneously to solve problems independently. You will define the future of our platform, driving the integration of AIOps to create self-healing, intelligent, and highly resilient systems.
If you are a visionary engineer who thrives on solving systemic challenges, optimizing performance at scale, and architecting autonomous infrastructure, this role is for you.
-
#### Key Responsibilities
**Platform Architecture & Strategy**
- Architect and evolve a unified, developer-centric platform across multi-cloud environments (AWS, Azure, and/or GCP).
- Bridge the gap between application logic and infrastructure, ensuring seamless integration and high-performance execution.
**AIOps & Intelligent Automation**
* **Lead the AIOps transformation:** Architect and implement AI-driven systems for predictive scaling, intelligent incident response, and automated root-cause analysis.
- Leverage LLMs and ML models to automate Infrastructure-as-Code (IaC) generation, security remediation, and complex diagnostic workflows.
- Implement log anomaly detection and predictive observability to transition from reactive to proactive system management.
**CI/CD & Developer Experience (DevEx)**
- Design high-velocity, "paved-road" deployment pipelines that empower engineers to ship code safely.
- Drive the adoption of AI-assisted development tools to accelerate the software delivery lifecycle.
**Kubernetes & Scalability**
- Lead the architecture of global-scale Kubernetes environments, focusing on multi-tenancy, security, and service mesh performance.
- Optimize system performance under extreme traffic, ensuring 99.99% availability through advanced capacity planning.
**FinOps & Governance**
- Implement advanced FinOps practices to optimize cloud spend, utilizing AI-driven forecasting to eliminate waste.
- Establish strong human-in-the-loop governance for all AI-automated infrastructure changes.
-
#### What We’re Looking For
**Experience & Expertise**
* **15+ years of experience** in Software Engineering, Infrastructure, and SRE/Platform Engineering.
* **Dual-Competency:** Proven track record of architecting large-scale applications *and* managing the underlying cloud infrastructure.
* **Autonomy:** Proven ability to troubleshoot issues across the entire stack—from application code and API performance to network latency and container orchestration.
**Technical & Problem-Solving Skills**
- Deep expertise in multi-cloud architecture and distributed systems design.
- Expert-level knowledge of Kubernetes, service meshes (e.g., Istio, Linkerd), and cloud-native observability.
* **AIOps Proficiency:** Demonstrated experience implementing AI/ML-driven operational workflows (e.g., automated diagnostics, predictive maintenance).
**Collaboration & Leadership**
- Experience leading and mentoring teams in a global, cross-functional environment.
- Ability to influence stakeholders and align technical strategy with high-level business objectives.
**Qualifications**
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
-
#### Nice to Have
- Experience training or fine-tuning models for infrastructure automation.
- Contributions to open-source projects related to Cloud-Native or AI tooling.
- Deep knowledge of modern observability stacks (Prometheus, Grafana, Honeycomb, Datadog).
- Experience with advanced infrastructure security (Zero Trust, Policy-as-Code).