





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, remote and metro role but highly specialized senior GPU infrastructure reduces applicant density.
Requires niche GPU/HPC, Slurm, InfiniBand and bare-metal skills that are not easily transferable across industries.
Explicit 8+ years, leadership and many mandatory niche skills (Slurm, NVidia bare-metal, InfiniBand) make filters strict.
Own end-to-end architecture and delivery of a bare-metal GPU infrastructure platform including Slurm scheduling layer and Kubernetes control plane.
Lead and line-manage a distributed engineering team (~12 members) across backend, frontend, DevOps, QA, and documentation for platform development and operations.
Act as primary technical interface to infrastructure partners and internal research/model-training teams, ensuring platform meets operational and capacity needs.
8+ years hands-on engineering experience with 3+ years leading infrastructure platform teams supporting other teams.
Hands-on experience running Slurm cluster for real users including controller, accounting, partitions, node health, and upgrades.
Strong expertise in bare-metal GPU fleet operations including NVIDIA driver/CUDA lifecycle, GPU health monitoring, and node onboarding.
Production Kubernetes operation experience including control plane lifecycle, operators, multi-tenancy and upgrades; JavaScript/Node.js fluency to review control plane code.
Technical leader comfortable in hands-on coding and architecture decisions, managing a distributed, multi-discipline engineering team across time zones.
Deep domain expertise in HPC/GPU infrastructure for research or AI workloads with experience in high-performance interconnects, storage, and multi-tenant GPU scheduling.
Experienced in partner/vendor technical collaboration, capacity planning, and managing platform delivery with a fixed timeline in a complex environment.