Job Purpose
The Linux & Kubernetes Administrator is responsible for the administration, maintenance, optimization, automation, and security of enterprise Linux platforms and containerized environments supporting mission-critical business applications. The role focuses on ensuring high availability, scalability, operational efficiency, and security compliance across cloud, on-premises and hybrid infrastructure environments leveraging Linux, Kubernetes, Docker, OpenShift, and cloud-native technologies
Roles & Responsibilities
- Linux Infrastructure Administration
Install, configure, administer, and support Red Hat Enterprise Linux (RHEL), Ubuntu, SUSE Linux, and other enterprise Linux distributions.
Manage system lifecycle activities, including provisioning, patching, upgrades, hardening, and decommissioning.
Monitor system health, availability, resource utilization, and performance metrics.
Manage user administration, access controls, and privilege management.
- Container Platform Administration
Deploy, administer, and support container platforms including Kubernetes, OpenShift, Docker Enterprise, Rancher, and container runtime environments.
Manage Kubernetes clusters across development, testing, and production environments.
Configure and maintain Namespaces, Pods, StatefulSets, Deployments, Services, Ingress Controllers, and Persistent Volumes.
Ensure container platform availability, scalability, and performance.
- Cloud-Native Platform Management
Support hybrid and multi-cloud containerized workloads running on L & T Vyoma, private cloud & Sovereign cloud environments.
Implement cloud-native operational best practices for containerized applications.
Manage container registries, image repositories, and artifact management solutions.
Support microservices-based application deployments.
Support AI/ML, GPU, and high-performance computing infrastructure.
Support GPU/ AI environment
- Platform Reliability & Performance Management
Identify and resolve Linux and container platform performance issues.
Conduct root cause analysis (RCA) for production incidents.
Perform proactive health checks, capacity planning, and resource optimization.
Implement monitoring and alerting for infrastructure and container workloads.
Implement Linux OS hardening as per CIS, STIG, and organizational standards.
Manage vulnerability remediation and security patching.
Configure security policies for Kubernetes and OpenShift environments.
Implement container image scanning, compliance monitoring, and runtime protection.
Ensure compliance with ISO 27001, PCI-DSS, SOC2, and organizational governance requirements.
- Automation & DevOps Enablement
Develop automation scripts using Shell Scripting, Python, Ansible, and PowerShell where applicable.
Implement Infrastructure as Code (IaC) using Terraform, Ansible, and GitOps methodologies.
Support CI/CD pipelines integrated with Kubernetes and container platforms.
Drive operational automation and service reliability improvements.
- Backup, Recovery & Business Continuity
Implement backup and recovery procedures for Linux and containerized platforms.
Support disaster recovery planning and testing.
Ensure platform resilience and adherence to RPO/RTO commitments.
Participate in business continuity and disaster recovery exercises.
- Monitoring & Observability
Administer enterprise monitoring solutions such as Prometheus, Grafana, ELK, Splunk, Dynatrace, AppDynamics, and Zabbix.
Configure dashboards, alerts, and operational reporting.
Analyze trends and recommend performance improvements.
Support observability initiatives across cloud-native platforms.
- Stakeholder & Service Management
Participate in Incident, Problem, Change, and Release Management activities.
Collaborate with Application, DevOps, Cloud Engineering, Security, and Infrastructure teams.
Support audits, compliance reviews, and customer governance meetings.
Mentor junior administrators and provide technical leadership.
Experience & Educational Requirement
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
RELEVANT EXPERIENCE
- Minimum 8-12 years of experience in Linux System Administration.
- Minimum 5+ years of hands-on experience supporting Kubernetes/OpenShift platforms in production environments.
- Experience managing large-scale enterprise and cloud-native environments.
- Strong exposure to container orchestration technologies.
- Experience supporting mission-critical 24x7 managed services operations.
- Experience in DevOps, CI/CD, Infrastructure as Code, and automation frameworks.
- Understanding of SRE principles and platform reliability engineering practices.
- Scripting knowledge in Python, Bash, or PowerShell.