Role:
Manage and optimize on-premises and cloud-based High-Performance Computing (HPC) infrastructure, CAE applications and engineering simulation environments to provide reliable, secure, scalable and cost-effective computing services for vehicle development and validation activities.
Responsibilities:
Administer and maintain High-Performance Computing (HPC) clusters, servers, storage, operating systems and network components.
Install, configure, monitor and troubleshoot HPC servers, job scheduling environment and associated infrastructure.
Administer and support CAE software, simulation applications and license management systems.
Provide user onboarding, access management and operational support for engineering users utilizing HPC resources.
Monitor system performance, utilization, queue status, job success rate and resource availability.
Support software deployment, upgrades, patch management, backup, security compliance and system health reporting.
Coordinate with IT teams, software vendors and engineering users for issue resolution and performance improvement. Support capacity planning, business continuity and disaster recovery activities for HPC infrastructure.
Maintain technical documentation, operational procedures, knowledge base and monthly utilization reports.
Contribute towards digital transformation initiatives related to simulation infrastructure and HPC modernization. Cloud HPC Server Management Responsibilities:
Provision, configure, administer and monitor cloud-based HPC servers, virtual machines, storage, networking and access controls.
Support hybrid HPC execution by coordinating workload movement between on-premises HPC and cloud HPC environments.
Monitor cloud resource consumption, prepare utilization/cost reports and identify opportunities for cost optimization.
Coordinate with cloud service providers and internal IT teams for service availability, incident resolution and performance troubleshooting.
Support cloud security controls, backup, disaster recovery, business continuity and compliance requirements for HPC workloads.
Evaluate cloud HPC services and contribute to the roadmap for scalable simulation computing infrastructure.
Technical Skills Preferred:
Linux Administration, HPC Cluster Management, Job Schedulers (SLURM/PBS), Server/Storage/Network Fundamentals, Bash/Shell/Python Scripting, FlexLM/FlexNet, CAE Software Exposure (Abaqus, LS-DYNA, ANSYS, HyperWorks, STAR-CCM+, Simcenter), Cloud HPC concepts (Azure/GCP/OCI), Monitoring, Backup and DR.