Hyderabad, Telangana
Job Summary
Seeking an experienced Support Engineer having experience of 7-10 years working as Site Reliability Engineer (SRE) with strong expertise in Kubernetes, Linux, Cloud Platforms (AWS/Azure/GCP), Observability, Automation, and Production Support . Responsible for managing and supporting production Kubernetes environments , ensuring platform reliability, availability, security, scalability, and operational excellence.
Key Responsibilities:
Manage and maintain Kubernetes clusters, including deployments, upgrades, capacity planning, RBAC, networking, storage, Ingress, ConfigMaps, Secrets, Services, Persistent Volumes, StatefulSets, DaemonSets, Jobs/CronJobs, Helm, and Autoscaling (HPA/VPA).
Provide 24x7 production support, incident management, RCA, postmortems, service restoration, and SLA/SLO compliance.
Monitor and improve platform reliability using Prometheus, Grafana, Loki, Elastic Stack, OpenTelemetry, and AlertManager .
Troubleshoot Kubernetes, Linux, container (Docker/OCI), networking (DNS, Load Balancers, TLS, Ingress), cloud, and infrastructure issues.
Automate operational tasks through Bash, Python, Terraform, and Ansible , and support Infrastructure as Code practices.
Support CI/CD and release management using GitHub Actions, GitLab CI, Jenkins, and ArgoCD (preferred) .
Perform patching, cluster maintenance, security updates, backups, disaster recovery validation, and platform upgrades.
Create runbooks, operational documentation, dashboards, alerts, and capacity planning reports.
Collaborate with Development, Platform Engineering, Security, Networking, Cloud Operations, and DevOps teams to improve system resilience and operational efficiency.
Required Skills: Kubernetes Administration, Linux, Docker/OCI, AWS/Azure/GCP, Networking, CI/CD, GitOps, Observability, Incident Management, RCA, Automation, Terraform, Ansible, Bash, Python.
Preferred: CKA/CKS certification, Cloud certifications, Multi-cluster/Multi-region Kubernetes, Service Mesh (Istio/Linkerd), High Availability, Disaster Recovery, Security Hardening, Capacity Planning, Performance Tuning, Cost Optimization, Chaos Engineering, AI-assisted Observability.
Key Competencies: Strong troubleshooting, ownership, production support, customer focus, communication, collaboration, continuous improvement, and ability to perform under pressure.
Success Metrics: High platform availability, improved MTTR, reduced incidents and alert noise, SLA/SLO compliance, increased automation coverage, successful upgrades/maintenance, and customer satisfaction.
Key Responsibilities
null
Other Requirements
Preferred Qualifications
Certified Kubernetes Administrator (CKA)
Certified Kubernetes Security Specialist (CKS)
Cloud certifications (AWS/Azure/GCP)
Experience supporting multi-cluster Kubernetes environments.
Experience with service mesh technologies (Istio/Linkerd).
Experience with GitOps workflows.
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-