Responsibilities
CORE RESPONSIBILITIES
- Define the platform SRE, DevSecOps, and security strategy - spanning architecture, toolchain selection, and
operational standards.
- Design for high availability, scalability, resilience, and disaster recovery across all platform tiers and deployment
regions.
- Establish observability, monitoring, logging, and incident management - own the full telemetry stack from
instrumentation to alerting to post-mortems.
- Govern security architecture, IAM, secrets management, and compliance - ensure the platform meets SOC 2, ISO
27001, and customer-specific security requirements.
- Define CI/CD pipelines, infrastructure automation (IaC), and operational runbooks; drive adoption of GitOps and policy as-code practices.
- Drive performance engineering and capacity planning; own SLOs, SLIs, and error budgets across all platform services.
- Partner with the Technical Architect and FDE leads to embed reliability and security earlier in the development lifecycle
- Lead incident response, blameless post-mortems, and reliability reviews; translate learnings into platform
improvements
Requirements
MUST-HAVE SKILLS & EXPERIENCE
- 12+ years in SRE, Platform Engineering, Cloud Infrastructure, or Cloud Security roles at enterprise or hyperscale scale.
- Expert-level Kubernetes: multi-cluster operations, operator patterns, network policies, pod security, and cluster
hardening.
- Deep cloud-native expertise across AWS, Azure, or GCP: VPC design, IAM, KMS, secrets management (Vault/AWS
Secrets Manager), and cloud security posture management.
- Strong DevSecOps practice: SAST/DAST tooling, container image scanning, SBOM, supply chain security, and secure
CI/CD pipeline design.
- Observability stack mastery: Prometheus, Grafana, OpenTelemetry, distributed tracing (Jaeger/Tempo), log
aggregation (ELK/Loki), and AIOps-ready alerting.
- Infrastructure-as-Code proficiency: Terraform, Pulumi, or CDK; GitOps with ArgoCD or Flux; policy-as-code with
OPA/Kyverno.
- Proven experience defining and operating SLOs, error budgets, and on-call practices in a high-stakes production
environment
Nice-to-have Skills
- Experience operating AI/ML platforms at scale - GPU cluster management, model serving infrastructure (vLLM, Triton),
and LLM observability.
- Compliance and audit experience: SOC 2 Type II, ISO 27001, GDPR, or telecom-specific security frameworks (NIST
Chaos engineering practice: Chaos Monkey, LitmusChaos, or equivalent - running Game Days and failure injection exercises.
- Telecom infrastructure context: network function virtualisation (NFV), cloud-native network functions (CNF), or 5G core platform operations.
MINDSET & BEHAVIOURS
- Reliability-first - treats every architectural decision through the lens of what fails, how it fails, and how fast you recover.
- Security-by-design - embeds security into the platform architecture, not bolted on post-deployment; champions zero trust principles.
- Data-driven operator - SLOs and error budgets are the language; makes reliability trade-offs with evidence, not intuition.
- Platform enabler - builds systems and tooling that make every other engineer faster and safer; infrastructure as a product.
Preferred Qualifications
- B.S. / M.S. in Computer Science, Systems Engineering, or equivalent; CKS (Kubernetes Security), AWS Security
Specialty, or CISSP a strong plus.
- Experience building and operating large-scale SaaS or enterprise products with 99.99%+ availability commitments.
- Track record leading SRE or platform functions in a telecom, fintech, or cloud-native software company
Pay: ₹634,401.85 - ₹2,079,151.76 per year
Work Location: In person