Role Overview
We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms . The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.
Key Responsibilities
- HPC Infrastructure Management
o Operate and maintain HPC clusters based on CentOS, RHEL , and hardware platforms like HPE and NVIDIA DGX .
o Ensure optimal performance, scalability, and reliability of compute resources.
o Manage large-scale storage systems including Dell Isilon , VAST Storage , Lustre , and GPFS .
o Implement data lifecycle management and optimize storage performance for HPC workloads.
o Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
o Troubleshoot network performance issues and ensure secure connectivity.
- Cluster and Job Scheduling
o Administer cluster management tools such as Bright Cluster Manager , Altair Grid Manager , and IBM LSF .
o Optimize job scheduling and resource allocation for diverse workloads.
- Monitoring and Automation
o Implement monitoring solutions using Zabbix , Grafana , and ELK Stack .
o Automate provisioning and configuration using Cobbler , Chef , Ansible , and AWS ParallelCluster .
- Performance Tuning & Troubleshooting
o Conduct performance benchmarking and tuning for HPC workloads.
o Diagnose and resolve hardware/software issues across compute, storage, and network layers.
o Ensure HPC environment adheres to security best practices and compliance standards.
Required Skills & Qualifications
o Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms ( HPE , NVIDIA DGX ).
o Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
o Proficiency in InfiniBand networking and high-speed interconnects.
o Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
o Expertise in Ansible , Chef , Cobbler , and scripting languages (Bash, Python).
o Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
o Practical experience with Zabbix , Grafana , and ELK Stack for system health and performance monitoring.
o Strong problem-solving and analytical skills.
o Ability to work in a fast-paced environment and lead technical teams.
o Excellent communication and documentation skills.
Preferred Qualifications
- Exposure to AI/ML workloads on HPC clusters.
- Experience with containerization (Docker, Singularity) in HPC environments.
- Knowledge of security hardening for HPC systems.
Education
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
#LI-LK1