





Tier-1 employer and mid-level requirement, but specialized HPC/GPU skills reduce broad competition.
Highly domain-specific HPC, GPU, InfiniBand, and Lustre expertise reduces cross-industry transferability.
Multiple mandatory technologies, explicit years, and specialized HPC/GPU experience required.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead systems administration, upgrades, incident response, and reliability improvements for large-scale AI/HPC clusters.
Own day-to-day operations ensuring system health, user satisfaction, efficient resource use, and cluster optimization.
Develop scalable automation and manage heterogeneous AI/ML clusters on-premises and cloud with a focus on user needs and workload performance.
Bachelor’s degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
Minimum 5 years experience designing and operating large-scale compute infrastructure.
Experience with AI/HPC job schedulers (e.g. Slurm, K8s, PBS, RTDA, BCM, LSF) and proficiency with CentOS/RHEL and/or Ubuntu Linux administration.
Skills in cluster configuration management tools (Terraform, Ansible, Puppet, Salt, BCM), container technologies (Docker, Singularity, Podman), Python and Bash scripting, and experience with AI/HPC workflows using MPI.
Experienced in managing distributed high-performance computing clusters with a strong technical leadership capacity on AI/HPC infrastructure.
Skilled in performance analysis and optimization of complex AI and HPC workloads with a user-focused approach to service and infrastructure reliability.
Familiarity with NVIDIA GPUs, CUDA, MLPerf benchmarking, AI/ML frameworks (PyTorch, Tensorflow), and high-speed networking/storage technologies (InfiniBand, Lustre, GPFS).