Applied AI Site Reliability Engineer III
Location: Hyderabad / Bengaluru
Experience: 5+ Years
Job Type: Full-time
About the Role
We are looking for an experienced Applied AI Site Reliability Engineer III to join a Product Engineering team focused on building and operating highly reliable, scalable, secure, and cost-effective cloud-native platforms and AI-enabled products.
The ideal candidate will have strong hands-on experience in Site Reliability Engineering, Cloud Platform Engineering, DevOps, Observability, Performance Engineering, Kubernetes, Infrastructure-as-Code, and AI/ML production operations.
Key Responsibilities
- Own and improve production reliability, performance, availability, scalability, and operational efficiency using SLIs, SLOs, SLAs, and error budgets.
- Design and implement production-grade observability using metrics, logs, tracing, dashboards, and actionable alerts.
- Drive production readiness, release validation, operational acceptance, and reliability checks.
- Manage cloud-native platforms and production environments across AWS, Azure, or GCP.
- Build and maintain infrastructure using Kubernetes, Docker, Terraform, CI/CD, and Git-based deployment practices.
- Perform load testing, performance testing, capacity planning, autoscaling, and resilience testing.
- Implement and support chaos engineering practices to validate system resilience.
- Participate in incident response, root-cause analysis, on-call operations, and blameless postmortems.
- Automate repetitive operational tasks and continuously reduce engineering toil.
- Operate AI/ML, GenAI, and agentic workloads in production, addressing reliability and performance challenges.
- Work with AI-specific production failure modes such as model drift, train/serve skew, output variance, latency, and token/GPU cost anomalies.
- Contribute to MLOps/LLMOps, AI control planes, model/LLM gateways, guardrails, and AI operational tooling.
- Implement cloud and AI cost optimization / FinOps practices.
- Collaborate with engineering, platform, security, risk, architecture, data governance, and product teams.
- Develop technical specifications, runbooks, operational playbooks, and reliability standards.
- Continuously improve system resilience, scalability, security, observability, and operational maturity.
Required Skills & Experience
- 5+ years of experience in Software Engineering, Site Reliability Engineering, DevOps, or Production Engineering.
- 3+ years of experience operating large-scale, distributed, cloud-native production systems.
- Hands-on experience with one or more programming/scripting languages such as Python, Go, Bash, Java, or C#/.NET.
- Strong experience with Kubernetes, Docker, Terraform, CI/CD, and cloud infrastructure.
- Experience with AWS, Azure, or GCP.
- Strong understanding of SLI, SLO, SLA, error budgets, incident management, production readiness, and on-call operations.
- Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations, or Splunk.
- Hands-on experience with performance/load testing using tools such as JMeter, LoadRunner, or k6.
- Experience with capacity planning, autoscaling, resilience testing, and chaos engineering.
- Experience operating AI/ML or GenAI workloads in production.
- Knowledge of MLOps, LLMOps, AI reliability, and agentic workloads.
- Experience with platforms such as Azure OpenAI, AWS Bedrock, or Vertex AI is highly desirable.
- Familiarity with MLflow, LangSmith, LangFuse, or equivalent AI/agent orchestration tools.
- Understanding of DevSecOps, SRE, Lean/XP methodologies, RBAC, secrets management, least privilege, and deployment controls.
- Strong software engineering fundamentals including OOP/OOD, data structures, algorithms, system design, and code instrumentation.
- Strong analytical, troubleshooting, communication, and stakeholder-management skills.
Preferred Candidate Profile
Candidates with experience in roles such as:
Senior Site Reliability Engineer | SRE | Production Engineer | Cloud SRE | Platform Engineer | DevOps Engineer | Reliability Engineer | AI/ML Platform Engineer | MLOps Engineer | LLMOps Engineer
Candidates who have worked on large-scale distributed systems, cloud platforms, SaaS products, AI/ML platforms, GenAI applications, or highly available enterprise applications will be preferred.
What We Offer
- Opportunity to work on cloud-native, AI-enabled enterprise platforms
- Exposure to GenAI, agentic AI, SRE, DevSecOps, and modern platform engineering
- Work with cross-functional engineering and architecture teams
- Opportunity to solve complex reliability, scalability, performance, and cost-engineering challenges
Travel: Up to 10% may be required.
Interested candidates can apply with their updated resume. Please mention your current CTC, expected CTC, notice period, and preferred location.
Work Location: In person