Experience: 6-8 years
Location: Pune
Role Overview:
We are looking for a Cloud & Infrastructure Engineer to design, implement, operate, and secure both on-premises Linux/Kubernetes infrastructure and public cloud environments used by ML Engineers, Infrastructure Engineers, and application teams. This role combines end-to-end Kubernetes cluster lifecycle management with cloud infrastructure engineering — spanning automation, infrastructure-as-code, networking, identity and access management, security hardening, CI/CD, and cost optimization.
Key Responsibilities:
AI-Enabled Engineering
Collaborate with Agentic AI models to optimize and streamline engineering processes (agentic optimization engineering)
Linux & Kubernetes Cluster Administration
Manage and operate Linux-based infrastructure, services, and systems
Own and operate Kubernetes clusters across their end-to-end lifecycle, including provisioning, configuration, maintenance, upgrades, and operational management
Support Kubernetes administration and container-platform operations
Cloud Infrastructure
Own and administer cloud environments and infrastructure for engineering teams
Design, implement, and operate scalable public, private, and hybrid cloud infrastructure
Design and implement cloud networking solutions
Monitor and optimize cloud infrastructure performance and availability
Perform cloud cost analysis, optimization, and ongoing cost management
Automation & Infrastructure as Code
Automate infrastructure and operational activities across on-prem and cloud environments
Build and manage Infrastructure as Code (IaC) using tools such as Terraform and Ansible
Support CI/CD pipelines and cloud-based application deployments
Security & Access Management
Apply infrastructure, platform, and cloud security controls
Implement and manage Identity and Access Management (IAM) and Role-Based Access Control (RBAC)
Manage permissions and ensure secure access to cloud and cluster resources
Support security hardening of operating systems and networks
Handle data-sensitivity and security-related configuration requirements
Collaboration
Work with ML Engineers, DevOps Engineers, and application teams to provision required infrastructure
Partner with engineering and infrastructure teams to maintain reliable enterprise environments
Must be available for on-call support as needed
Ability to collaborate effectively with Agentic AI models
Capability to optimize repetitive processes and workflows using agentic/autonomous engineering approaches
Required Skills & Experience:
Technical Skills
Strong Linux administration skills
Strong Kubernetes administration and cluster lifecycle management experience; Helm knowledge preferred
Strong experience with at least one major public cloud platform (AWS, Microsoft Azure, or GCP); exposure to multiple platforms preferred
Experience working across public, private, and hybrid cloud environments
Hands-on experience with Infrastructure as Code — Terraform and/or Ansible (or equivalent)
Strong knowledge of cloud networking, IAM, RBAC, and security controls
Experience with CI/CD pipelines and deployment processes
Understanding of OS and network hardening
Understanding of infrastructure security and data-sensitivity controls
Soft Skills
Strong troubleshooting and analytical skills
Strong sense of ownership over Kubernetes and cloud environments
Ability to work effectively with cross-functional engineering teams
Good communication and documentation skills
Certifications (Preferred)
Relevant Kubernetes certifications such as CKA/CKAD
Professional-level cloud certifications in AWS, Azure, or GCP (entry-level/foundational certifications carry less weight)