Overview:
ROLE PURPOSE
The SRE Architect owns the reliability, security, and operational excellence of the Autonomous Operations platform. This
role defines and enforces the engineering discipline that keeps Prodapt's AI-native products running at enterprise scale -
designing systems for high availability and resilience, establishing the observability fabric that makes failures visible, and
governing the security posture that makes customers trust the platform. They are the last line of defence between a great
product and a production incident.
Responsibilities:
CORE RESPONSIBILITIES
- Define the platform SRE, DevSecOps, and security strategy - spanning architecture, toolchain selection, and
operational standards.
- Design for high availability, scalability, resilience, and disaster recovery across all platform tiers and deployment
regions.
- Establish observability, monitoring, logging, and incident management - own the full telemetry stack from
instrumentation to alerting to post-mortems.
- Govern security architecture, IAM, secrets management, and compliance - ensure the platform meets SOC 2, ISO
27001, and customer-specific security requirements.
- Define CI/CD pipelines, infrastructure automation (IaC), and operational runbooks; drive adoption of GitOps and policyas-code practices.
- Drive performance engineering and capacity planning; own SLOs, SLIs, and error budgets across all platform services.
- Partner with the Technical Architect and FDE leads to embed reliability and security earlier in the development lifecycle
(shift-left).
- Lead incident response, blameless post-mortems, and reliability reviews; translate learnings into platform
improvements
Requirements:
MUST-HAVE SKILLS & EXPERIENCE
- 12+ years in SRE, Platform Engineering, Cloud Infrastructure, or Cloud Security roles at enterprise or hyperscale scale.
- Expert-level Kubernetes: multi-cluster operations, operator patterns, network policies, pod security, and cluster
hardening.
- Deep cloud-native expertise across AWS, Azure, or GCP: VPC design, IAM, KMS, secrets management (Vault/AWS
Secrets Manager), and cloud security posture management.
- Strong DevSecOps practice: SAST/DAST tooling, container image scanning, SBOM, supply chain security, and secure
CI/CD pipeline design.
- Observability stack mastery: Prometheus, Grafana, OpenTelemetry, distributed tracing (Jaeger/Tempo), log
aggregation (ELK/Loki), and AIOps-ready alerting.
- Infrastructure-as-Code proficiency: Terraform, Pulumi, or CDK; GitOps with ArgoCD or Flux; policy-as-code with
OPA/Kyverno.
- Proven experience defining and operating SLOs, error budgets, and on-call practices in a high-stakes production
environment
NICE-TO-HAVE SKILLS
- Experience operating AI/ML platforms at scale - GPU cluster management, model serving infrastructure (vLLM, Triton),
and LLM observability.
- Compliance and audit experience: SOC 2 Type II, ISO 27001, GDPR, or telecom-specific security frameworks (NIST
CSF, ENISA).
Hiring Profiles | AI-Native Product Team | 2026
Prodapt Solutions Pvt. Ltd. | Confidential Page 2 of 2
- Chaos engineering practice: Chaos Monkey, LitmusChaos, or equivalent - running Game Days and failure injection
exercises.
- Telecom infrastructure context: network function virtualisation (NFV), cloud-native network functions (CNF), or 5G core
platform operations.
MINDSET & BEHAVIOURS
- Reliability-first - treats every architectural decision through the lens of what fails, how it fails, and how fast you recover.
- Security-by-design - embeds security into the platform architecture, not bolted on post-deployment; champions zerotrust principles.
- Data-driven operator - SLOs and error budgets are the language; makes reliability trade-offs with evidence, not intuition.
- Platform enabler - builds systems and tooling that make every other engineer faster and safer; infrastructure as a
product.
PREFERRED QUALIFICATIONS
- B.S. / M.S. in Computer Science, Systems Engineering, or equivalent; CKS (Kubernetes Security), AWS Security
Specialty, or CISSP a strong plus.
- Experience building and operating large-scale SaaS or enterprise products with 99.99%+ availability commitments.
- Track record leading SRE or platform functions in a telecom, fintech, or cloud-native software company