Focus: Traditional infrastructure automation + cloud fluency
About the Role
Join a team responsible for running reliable, scalable infrastructure operations day-to-day — owning observability, reliability, and vulnerability patching across on-prem and cloud environments, rather than one-off build projects.
Responsibilities
Own observability and monitoring coverage (Nagios, Graphite, Zabbix, etc.) across data center and cloud infrastructure — dashboards, alerting, and health tracking.
Drive day-to-day reliability operations: incident response, troubleshooting, on-call support, and proactive maintenance across hybrid infrastructure.
Own vulnerability patching cadence across on-prem Linux systems and AWS/Azure/GCP cloud resources, tracking issues to closure.
Automate recurring operational tasks (Shell/Python/Ruby/Golang) — patching, health checks, monitoring responses — to reduce manual toil.
Maintain configuration-as-code (Puppet manifests, Terraform/CloudFormation) as part of ongoing operational upkeep.
Collaborate across infrastructure teams on recurring reliability/security issues and contribute operational improvements.
Requirements
Strong Linux systems administration and internals knowledge, with solid troubleshooting skills.
Scripting/programming proficiency: Shell, Python, Perl, or Golang.
Strong networking fundamentals: TCP/IP, switches, routers, load balancers.
Hands-on experience with observability/monitoring tools (Nagios, Graphite, Zabbix, etc.) and vulnerability patching workflows.
Experience with Puppet, Chef, Ansible, or similar orchestration tools.
Working knowledge of at least one public cloud platform (AWS preferred; GCP/Azure a plus) and cloud/IAC basics.
Pay: ₹638,651.49 - ₹1,100,000.36 per year
Work Location: In person