About the Role
We are looking for a highly skilled and motivated Senior DevOps Engineer to join our platform engineering team. In this role, you will be a key contributor to the design, delivery, and evolution of our cloud infrastructure and Kubernetes platforms. You are someone who thrives in ambiguity, brings a growth mindset to every challenge, and communicates clearly across engineering and business stakeholders.
Key Responsibilities
-
Design, build, and maintain scalable, secure, and highly available infrastructure across multi-account, multi-region, and multi-AZ AWS environments
-
Lead Infrastructure as Code practices using Terraform — authoring reusable modules, managing workspaces, and enforcing standards across teams
-
Administer and evolve Amazon EKS clusters including upgrades, secondary networking configurations, and cluster lifecycle management
-
Write and maintain Helm charts including umbrella charts and integration of OSS/community charts; troubleshoot complex Helm deployment issues
-
Develop and maintain custom Kubernetes controllers and operators using standard tooling
-
Define and champion multi-cloud architecture patterns and best practices
-
Analyze open-ended infrastructure and platform problems, propose well-reasoned solutions, and drive them to completion with minimal supervision
-
Collaborate closely with development, security, and architecture teams — communicating complex technical concepts clearly to diverse audiences
-
Mentor junior engineers and contribute to a culture of continuous learning and improvement
Required Skills & Experience
Core Competencies
-
Strong communication skills — written and verbal — with the ability to present technical proposals to both engineers and non-technical stakeholders
-
Proven ability to analyze ambiguous, open-ended problems and independently propose and deliver pragmatic solutions
-
Strong problem-solving mindset with a structured, methodical approach to debugging and root-cause analysis
-
Demonstrated growth mindset — actively learns from incidents, seeks feedback, and keeps up with evolving industry practices
Terraform
-
hands-on experience writing production-grade Terraform configurations and reusable modules
-
experience with Terraform workspaces for environment and account segregation
-
experience using Terraform providers to provision and manage AWS resources across multi-account, multi-region, and multi-AZ topologies
-
experience provisioning and managing EKS clusters, VPC networking, and related AWS services via HashiCorp Terraform modules
-
experience configuring EKS secondary networking (e.g., VPC CNI custom networking) via Terraform
-
Good understanding of multi-cloud infrastructure architecture and portability considerations
Kubernetes & EKS
-
hands-on experience in Amazon EKS administration — cluster provisioning, IAM integration, node group management, and day-2 operations
-
experience as a Kubernetes practitioner — workload management, RBAC, networking, storage, autoscaling, and observability
-
understanding of Custom Resource Definitions (CRDs) and consumption of OSS Helm charts and Kubernetes operators
-
experience managing and executing EKS cluster version upgrades with minimal disruption
-
Hands-on experience writing custom controllers and operators using frameworks such as controller-runtime or Kubebuilder
-
hands-on experience with AWS Load Balancer Controller — ingress and service configuration on EKS
-
hands-on experience with Kubernetes Gateway API — HTTPRoute, GatewayClass, and migration from Ingress
Helm
-
experience authoring Helm charts for complex, production workloads
-
experience building and managing umbrella charts for multi-component application deployments
-
understanding of CRDs and OSS Helm charts — including evaluation, customisation, and maintenance
-
experience troubleshooting Helm release failures, hook issues, and chart rendering problems
Istio
-
hands-on experience deploying and operating Istio in production Kubernetes environments
-
experience with both Istio Sidecar (proxy) and Ambient mesh deployment modes
-
understanding of Istio core concepts and traffic management — VirtualService, DestinationRule, Gateway, AuthorizationPolicy, mTLS, and observability integration
AWS
-
hands-on experience designing and operating multi-region, multi-AZ AWS architectures
-
Deep working knowledge across core AWS services, including:
-
EKS — cluster administration, networking, and integrations
-
VPC — subnets, route tables, security groups, VPC peering, Transit Gateway
-
EC2 — instance types, launch templates, Auto Scaling Groups
-
S3 — bucket policies, lifecycle rules, replication
-
AWS API Gateway
-
Networking — NAT, IGW, PrivateLink, VPN, Direct Connect concepts
-
ACM (AWS Certificate Manager) — certificate provisioning, renewal, and integration with AWS services
-
PCA (AWS Private Certificate Authority) — private CA management for internal TLS
-
Route 53 — DNS management, health checks, routing policies
-
Strong understanding of TLS — certificate chains, mTLS, termination strategies, and PKI fundamentals
Nice to Have
-
Experience with GitOps tooling (e.g., ArgoCD, GitHub Actions, Cloudbees Jenkins)
-
Familiarity with observability stacks (e.g., Prometheus, Grafana, OpenTelemetry)
-
Experience with CI/CD pipelines (e.g., GitHub Actions, Jenkins, Atlantis)
Exposure to service mesh technologies (e.g., Istio, Linkerd)