Overview:
We are looking for a Site Reliability Engineer who can design, build, and operate resilient infrastructure and CI/CD pipelines across multiple cloud platforms (AWS, Azure, GCP). This role is central to enabling reliable, scalable, and secure deployment of applications and AI/data platforms, with a strong focus on automation, observability, and infrastructure-as-code across heterogeneous cloud environments.
Role Title: Site Reliability Engineer (SRE) – Multi-Cloud Infrastructure & CI/CD
Domain: Cloud Infrastructure, DevOps/SRE, Platform Engineering
Responsibilities:
-
Design, build, and maintain cloud-agnostic infrastructure using Infrastructure-as-Code (Terraform, Pulumi, or equivalent) across AWS, Azure, and GCP
-
Build and manage CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Azure DevOps, Cloud Build) supporting multi-cloud deployments
-
Define and implement SRE practices: SLIs/SLOs/SLAs, error budgets, incident response, and on-call processes
-
Set up monitoring, logging, and observability stacks (Prometheus, Grafana, Datadog, Cloud Monitoring/Stackdriver, ELK/EFK) across environments
-
Architect and manage container orchestration platforms (Kubernetes — EKS/AKS/GKE) and containerization (Docker) for portable, cloud-agnostic workloads
-
Automate provisioning, scaling, patching, and configuration management (Ansible, Chef, or Puppet as needed)
-
Implement disaster recovery, backup, high-availability, and multi-region/multi-cloud failover strategies
-
Drive cost optimization and capacity planning across cloud providers
-
Build self-healing systems and automate incident remediation to reduce toil
-
Conduct root cause analysis (RCA) and post-incident reviews; drive reliability improvements
-
Implement security best practices: IAM policies, secrets management (Vault, Cloud KMS), network security, and compliance guardrails across clouds
-
Collaborate with development teams to embed reliability, observability, and deployment best practices into the software delivery lifecycle (DevSecOps/shift-left)
-
Support migration or workload portability initiatives between cloud providers or hybrid/on-prem environments
-
Requirements:
-
5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles
-
Proven hands-on experience across at least two of the three major cloud providers (AWS, Azure, GCP) — true multi-cloud experience strongly preferred
-
Strong expertise in Infrastructure-as-Code (Terraform mandatory; CloudFormation/Bicep/Pulumi a plus)
-
Deep knowledge of Kubernetes and container orchestration in production environments
-
Strong scripting/programming skills (Python, Go, or Bash)
-
Experience building and maintaining CI/CD pipelines end-to-end (build, test, deploy, rollback)
-
Solid understanding of networking fundamentals (VPCs, load balancers, DNS, service mesh — Istio/Linkerd a plus)
-
Experience with observability and monitoring tooling, and setting up alerting that ties to meaningful SLOs
-
Familiarity with GitOps practices (ArgoCD, FluxCD)
-
Strong incident management experience — triage, escalation, blameless postmortems
-
Understanding of security and compliance requirements in cloud environments (encryption, IAM, network segmentation)
-
Cloud certifications across multiple providers (e.g., AWS Solutions Architect/SysOps, Azure Administrator/DevOps Engineer, GCP Professional Cloud DevOps Engineer)
-
Certified Kubernetes Administrator (CKA) or CKAD
-
Experience with service mesh, API gateways, and edge/CDN configuration
-
Experience supporting AI/ML or data platform infrastructure (GPU provisioning, vector databases, model serving infra)
-
Prior experience in telecom, enterprise SaaS, or large-scale regulated environments
-
Familiarity with FinOps practices for multi-cloud cost governance