





Tier-1 employer, metro location, and mid-level seniority increase applicant competition despite niche GPU cluster skills.
Requires specialized GPU cluster, InfiniBand, and Slurm experience reducing cross-industry transferability.
Explicit 5+ years, mandatory cluster administration, Ansible, Linux and networking make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain automation tools for deploying and managing large GPU clusters interconnected via NVLink and InfiniBand.
Own daily troubleshooting and resolution of cluster failures to ensure high availability and performance.
Manage software and firmware rollouts/rollbacks on clusters and coordinate with cross-functional engineering teams across time zones.
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
5+ years of hands-on experience deploying and administering clusters, servers, switches, and related infrastructure.
Strong skills in automation with Ansible, Python, and Shell scripting.
Proficiency with Linux fundamentals and deep understanding of operating systems, networks, and high-performance applications.
Experience with resource scheduling managers, preferably Slurm, and familiarity with industry-standard alerting tools and emergency response.
Hands-on experience working with GPU-focused hardware/software such as DGX systems and Compute Clusters.
Ability to design and implement large-scale networking solutions and metrics collection and alerting infrastructures.