





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, remote mid-level generalist role increases applicant density despite Slurm niche.
Requires niche Slurm and HPC experience, limiting cross-industry transferability.
Explicit 5+ years Slurm production experience and specialized HPC infrastructure skills make shortlisting stringent.
Own end-to-end support cases for production AI and HPC clusters using Slurm, from initial investigation through resolution.
Diagnose and resolve complex issues across Slurm components (slurmctld, slurmd, slurmdbd), scheduling, node/resource management, and integrated systems like Linux, networking, storage, authentication, and GPUs.
Advise customers on configuration, upgrades, incident recovery, and collaborate with engineering by producing technical reports and developing support knowledge resources.
5+ years hands-on experience administering and supporting Slurm in production HPC or AI environments including managing outage incidents.
BS degree in Computer Science, Engineering, or related field, or equivalent practical experience.
Expert-level understanding of Slurm architecture, daemons, configuration, scheduling, accounting, and failure modes.
Strong Linux system administration experience, including systemd, cgroups, authentication, networking, and database-backed services.
Experienced in operating Slurm on large-scale, multi-user clusters with complex scheduling policies and heterogeneous compute resources.
Capable of independently diagnosing sophisticated Slurm incidents and guiding them to technically sound resolutions.
Familiarity with Slurm internals, source code, plugins, container technologies, or integration with cluster management platforms is a strong plus.