





Strong VC-backed brand, metro location, mid-level (5+ years) role, but niche HPC specialization reduces broad applicant volume.
Highly specialized HPC GPU SRE skills transfer poorly across non-AI industries.
Explicit 5+ years, 2+ GPU-cluster requirement, on-call ownership, and mandatory Python/Go skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Manage and operate a large multi-vendor GPU fleet for demanding workloads including extended GPU training jobs and latency-sensitive inference services.
Own end-to-end fleet operations including provisioning, observability, capacity planning, and health monitoring with meaningful on-call responsibilities including incident response and postmortems.
Develop internal tools to support fleet operations and collaborate closely with ML and platform teams to ensure workload stability and performance.
5+ years in infrastructure or site reliability engineering with at least 2 years operating GPU clusters at scale (exceptions for Storage and Fabric experts with deep domain expertise).
Proficiency in Python or Go for building and maintaining internal tooling.
Demonstrated on-call infrastructure ownership with a track record of effective postmortems leading to durable fixes.
Working fluency across multiple technical areas related to GPU fleet operation including storage, network fabric, kernel/driver issues, distributed training faults.
Specialist with deep expertise in one of five core focus areas (not explicitly defined here) and good cross-team technical fluency to triage multi-domain GPU fleet problems.
Experienced in hybrid Slurm and Kubernetes environments, on-prem GPU deployments, and coordination with datacenter operations is a plus.
Comfortable operating in a high-ownership SRE role that requires technical depth, cross-functional partnership, and delivering high reliability for complex AI infrastructure.