Job Title
Kubernetes Platform Engineer
Role Summary
The Kubernetes Platform Engineer is responsible for designing, building, operating, and continuously improving critical, production‑grade Kubernetes platforms. This role focuses on platform reliability, security, scalability, and developer enablement across mission‑critical workloads.
Key Responsibilities
Design, deploy, and operate high‑availability Kubernetes clusters supporting business‑critical applications
-
Build and maintain platform capabilities including:
o Cluster lifecycle management (install, upgrade, patching)
-
Networking, ingress/egress, service mesh, and DNS
-
Storage (CSI), persistence, and backup/recovery
-
Implement security‑by‑default controls:
o RBAC, pod security, network policies, secrets management
-
Image scanning, runtime security, and policy enforcement
-
Establish observability standards (metrics, logs, traces) and SLO‑driven reliability practices
-
Automate platform operations using Infrastructure as Code (IaC)
-
Partner with application teams to:
o Enable self‑service consumption
-
Troubleshoot complex production issues
-
Define best practices for deployment and runtime behavior
-
Support incident response, root cause analysis, and continuous resilience improvement
-
Contribute to platform standards, reference architectures, and technical documentation
Required Qualifications
- Strong hands‑on experience operating production Kubernetes at scale
-
Experience with enterprise platforms such as OpenShift, EKS, AKS, or GKE
-
Deep understanding of Kubernetes internals:
o Scheduling, controllers, networking, and API mechanics
-
Experience with IaC tools (Terraform, Helm)
-
Experience designing for security, reliability, and fault tolerance
-
Proficiency with Python scripting language
-
Familiarity with CI/CD platform integration
-
Prior on‑call or operational ownership of critical services
-
Detail oriented, self-directed, Agile team player
What Success Looks Like
- Kubernetes platforms are stable, secure, and scalable
-
Application teams deploy faster with fewer production issues
-
Platform upgrades and changes are predictable and low‑risk
-
Clear guardrails exist without slowing innovation
-
Assignments are fulfilled on-time, in-scope and with a minimal degree of managerial oversight