Role description
Cloud Operations Engineer
Job Summary
We are looking for a skilled Cloud Operations Engineer to manage, monitor, and optimize our cloud infrastructure across AWS/Azure/GCP environments.
The role involves ensuring high availability, performance, and cost-efficiency of cloud workloads while supporting deployment pipelines and incident response, working from our Pune office.
Key Responsibilities
- Monitor cloud infrastructure (compute, storage, networking, databases) for availability, performance, and capacity across production and non-production environments.
• Manage day-to-day cloud operations including provisioning, scaling, patching, and decommissioning of cloud resources.
• Respond to and resolve infrastructure incidents and s within defined SLAs; perform root cause analysis (RCA) for recurring issues.
- Automate routine operational tasks using scripting (Bash, Python, PowerShell) and Infrastructure-as-Code tools (Terraform, CloudFormation, ARM templates).
- Manage backup, disaster recovery, and business continuity processes for cloud workloads.
- Implement and maintain cost optimization practices — rightsizing, reserved instances, auto-scaling policies.
- Support CI/CD pipelines and work with DevOps teams to deploy applications and infrastructure changes.
- Maintain cloud governance — tagging standards, access controls (IAM), and compliance with internal policies.
- Manage cloud monitoring/observability tools (CloudWatch, Azure Monitor, Datadog, Grafana, Prometheus).
- Participate in on-call rotation for critical incident support.
- Document runbooks, SOPs, and architecture diagrams for cloud environments.
- Coordinate with vendors/cloud providers for support tickets and escalations.
Required Skills & Qualifications
- Bachelor’s degree in computer science, IT, or related field.
- Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP); multi-cloud exposure is a plus.
• Working knowledge of Infrastructure-as-Code (Terraform/CloudFormation) and configuration management (Ansible/Chef/Puppet).
• Proficiency in scripting languages (Python, Bash, PowerShell).
• Understanding of networking fundamentals (VPC, subnets, load balancers, DNS, VPN).
- Familiarity with monitoring/logging tools.
- Relevant certifications preferred: AWS Certified SysOps Administrator, Azure Administrator Associate, or equivalent.
- Strong troubleshooting, analytical, and communication skills.
Skills
High Availability and Disaster Recovery, Infrastructure as Code, Google Cloud Platform, PowerShell
About UST
UST is a global digital transformation solutions provider. For more than 20 years, UST has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, UST partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, UST embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, UST builds for boundless impact—touching billions of lives in the process.