Location: Chennai
Experience: 5-8+ years
Open Positions: 01
Role Summary
Lead reliability engineering for business-critical manufacturing platforms, owning SLOs, observability, CI/CD, infrastructure automation, incident management and production readiness. Lead AI/Agentic-AI reliability covering model/agent observability, AIOps, drift monitoring, workflow guardrails and GenAI-assisted incident response.
Key Responsibilities
- Define SLI/SLO/SLA and error budgets; align release velocity with reliability targets.
- Build observability across metrics, logs and traces using Splunk, Datadog, Prometheus and Grafana.
- Create actionable alerts and dashboards that reduce alert fatigue.
- Develop Jenkins CI/CD pipelines and evaluate GitOps/Argo CD.
- Manage GCP infrastructure with Terraform; operate Docker/Kubernetes workloads.
- Conduct load/stress testing and production-readiness reviews.
- Own incident response, on-call, MTTA/MTTR, postmortems, runbooks and improvement.
- Monitor model-serving, agent orchestration and RAG systems for availability, latency, task success and drift.
- Define AI SLOs and safeguards including auditing, human-in-the-loop escalation and runaway-agent circuit breakers.
- Implement AIOps for anomaly detection, predictive alerts, automated triage and GenAI-assisted RCA/runbooks.
- Reduce operational toil through automation; collaborate across platform, manufacturing, data/AI and IT teams; mentor engineers.
- Provide weekly reliability, incident and KPI reporting.
Required Qualifications
- 5–8+ years in SRE, DevOps or platform engineering.
- Strong SLI/SLO, error-budget, observability and incident-management expertise.
- Strong GCP experience including networking, compute, managed services and access.
- Production experience with Kubernetes, Docker and Terraform.
- Jenkins CI/CD experience; GitOps/Argo CD exposure desirable.
- Splunk plus Datadog, Prometheus or Grafana experience.
- Python and/or Java scripting capability.
- Load/stress testing, capacity analysis and production-readiness experience.
- PagerDuty, Opsgenie or equivalent on-call tooling knowledge.
- Exposure to AI/ML reliability, MLOps, AIOps or Agentic-AI operations.
AI & Agentic-AI Reliability Experience
- Vertex AI: model monitoring, pipelines and agent-building services.
- LangChain, LangGraph or comparable agent orchestration frameworks.
- RAG, vector search and retrieval-system observability.
- Model performance and data-drift monitoring.
- Inference latency, model availability and agent task-completion observability.
- GenAI-assisted incident response, runbook retrieval, incident copilots and RCA summarisation.
Behavioural & Leadership Competencies
- Technical leadership and architectural decision-making.
- Analytical troubleshooting of distributed and AI-system failures.
- Clear communication of reliability, risks and progress to technical/executive audiences.
- Cross-functional collaboration and SRE mentoring.
- Operational ownership through incident closure, corrective actions and documentation.
Key Deliverables
- SLO/error-budget framework and integrated observability dashboards.
- Controlled CI/CD and Terraform-based infrastructure automation.
- Incident-management process, on-call model, runbooks and postmortems.
- Load/stress testing evidence for production readiness.
- Initial AI/Agentic-AI observability and AIOps/GenAI incident-response proof of concept.
- Modular onboarding architecture and regular reliability/KPI reporting.
Initial Success Measures
- SLOs operational for at least three critical platforms within three months.
- Initial model-monitoring or AIOps capability operational within three months.
- Critical-incident MTTA below 15 minutes with continuous MTTR improvement.
- Operational toil at or below 50% per sprint, with remaining capacity for automation.
Additional Expectations
- Flexibility for onsite collaboration, milestones and on-call participation.
- Approved engineering workstation with GCP and AI/ML tooling access.
- Strong documentation and modular architecture to support handover and scale-up.
Send your CVs to [email protected]