Infra and SRE (1987) Pune · Full Time · Hybrid
About the Role Manage and maintain large-scale hybrid infrastructure environments to ensure high availability, resiliency, scalability, and performance of critical business applications and infrastructure services. Lead proactive monitoring, incident response, and infrastructure health assessments. Drive service availability improvements through automation and operational excellence initiatives.
Key Responsibilities
- Manage large-scale hybrid infrastructure environments across on-premises and cloud platforms.
- Ensure high availability, resiliency, scalability, and performance of critical business applications and infrastructure services.
- Lead proactive monitoring, incident response, problem management, and infrastructure health assessments.
- Perform root cause analysis for major incidents and implement preventative measures.
- Drive service availability improvements through automation and operational excellence initiatives.
- Support and optimize Azure, AWS, Azure Local, and hybrid cloud environments.
- Design and implement resilient infrastructure architectures supporting business continuity requirements.
- Collaborate with application, cloud, security, and infrastructure teams during cloud adoption programs.
- Lead DC/DR planning, execution, testing, and governance activities.
- Implement infrastructure provisioning and configuration management using Terraform and Ansible.
Requirements
- Experience in managing large-scale hybrid infrastructure environments.
- Knowledge of Azure, AWS, Azure Local, and hybrid cloud environments.
- Understanding of disaster recovery and business continuity principles.
- Experience with Terraform and Ansible for infrastructure provisioning and configuration management.
- Strong problem-solving skills for root cause analysis and incident response.
What We Offer
- Opportunity to work with large-scale hybrid infrastructure environments.
- Collaboration with cross-functional teams for cloud adoption programs.
- Professional development through automation and operational excellence initiatives.