Role description
Job Code 14217 Locations Bengaluru Minimum Experience 6 Maximum Experience 10 Mandatory Skills Python, Terraform, Kubernetes, Ci/Cd, Site Reliability Engineering, AWS Experience 6 to 10 Years Location Bengaluru Job Description Reliability & Operations - Design, build, and maintain scalable platform services and internal developer tooling using Python - Develop automation frameworks, orchestration systems, and backend infrastructure services - Improve platform reliability, observability, deployment automation, and operational excellence - Drive infrastructure-as-code and cloud automation initiatives - Implement best practices around security, monitoring, testing, and production operations - Design, implement, and maintain highly available and resilient systems in Kubernetes-based environments - Contribute to architectural decisions for distributed systems and cloud-native applications - Mentor engineers and promote engineering excellence across the organization - Lead incident response, RCA, and postmortems - Drive reliability improvements through automation Cloud & Platform Engineering - Build and manage infrastructure on AWS. - Operate Kubernetes clusters (EKS preferred) - Deploy services using Helm, ArgoCD and Argo rollout - Manage containerized workloads using Docker and containerd - Terraform - Python Programming - Ansible - Packer - GitHub Actions / Jenkins / ArgoCD - Prometheus / Grafana / Datadog/ Splunk Automation & Tooling - Strong Python skills with emphasis on reliability, automation, and observability tooling - Develop automation and tooling using Python - Create internal reliability and monitoring tools - Integrate CI/CD pipelines with observability and reliability checks Collaboration & Leadership - Mentor junior engineers - Influence architecture decisions - Collaborate across engineering teams Required Qualifications - 6+ years of relevant experience in SRE, DevOps, or Platform Engineering - Strong Python skills with experience building production-grade automation and tooling - Strong programming experience in Python - Production experience with Kubernetes - Strong observability fundamentals - Experience with Helm, ArgoCD, Argo Rollout and Docker, Ansible, Packer - Experience with AWS cloud - Strong Linux and networking fundamentals - Familiarity with the SDLC Preferred Qualifications - Experience with tool chain development using Python - AWS Experience - Ansible, Packer - Multi-cluster or multi-region Kubernetes experience - Experience with Kubernetes package manager (helm) and deployment (ArgoCD / Argo Rollout) - Service mesh (Istio) and API gateway (Kong) experience - Infrastructure-as-Code (Terraform preferred) - Cloud cost optimization experience Project Details / What You’ll Work On As a Senior Platform Engineer, you will build and scale the core platform systems that enable engineers to develop, deploy, and operate services efficiently. Your primary focus will be designing Python-based platform services, automation frameworks, and cloud infrastructure that improve developer productivity and system reliability. You’ll Work On Building internal platform tools and APIs using Python Developing automation for infrastructure provisioning, deployments, and operational workflows Designing scalable backend services and platform components used across engineering teams Improving CI/CD pipelines, deployment reliability, and developer experience Managing and optimizing cloud-native infrastructure on Kubernetes and public cloud platforms Building observability solutions including monitoring, logging, ing, and tracing Reducing operational toil through automation and self-service engineering platforms Partnering with product engineering, SRE, and security teams to improve platform scalability and reliability Driving best practices around system design, infrastructure-as-code, testing, and production operations Education Qualificaiton B.E. Job Title Sr. Platform Engineer (Python) Roles & Responsibilities About the Role We are looking for a Senior Platform Engineer with good Python expertise to help build and scale the core platform that powers our engineering ecosystem. You will design and operate reliable backend services, developer tooling, infrastructure automation, and cloud-native systems that enable product teams to move fast safely and efficiently. This role sits at the intersection of software engineering, platform engineering, cloud infrastructure, and developer experience. You’ll work closely with product engineers, SREs, security teams, and architecture leadership to create scalable internal platforms and operational excellence across the organization. If you enjoy solving complex distributed systems problems, improving engineering productivity, and building robust Python-based platform services, we’d love to talk to you. Project Details Project Details / What You’ll Work On Build and operate a centralized observability platform for metrics, logs, traces, and ing across Kubernetes workloads using Prometheus, Grafana, OpenTelemetry, and GCP Cloud Monitoring Define and drive SLOs, SLIs, and error budgets to improve reliability, reduce MTTR, and guide release decisions Design, operate, and optimize EKS/GKE-based Kubernetes platforms using Helm and containerized workloads with Docker Develop Python-based automation and tooling for observability, SLO reporting, incident response, and operational workflows Lead incident response for production issues, conduct blameless postmortems, and drive long-term reliability improvements Optimize platform scalability, performance, and cloud cost efficiency with a strong focus on GCP and AWS. Act as a technical leader, influencing architecture and mentoring teams on reliability and observability best practices
Skills
Python, Kubernetes, Terraform, CI/CD
About UST
UST is a global digital transformation solutions provider. For more than 20 years, UST has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, UST partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, UST embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, UST builds for boundless impact—touching billions of lives in the process.