





VC-backed brand plus mid-level SRE title increases competition despite niche GPU-cluster specialization.
Highly specialized HPC SRE skills limit cross-industry transferability.
Explicit 5+ years, 2+ years GPU-cluster requirement, and tooling/on-call mandates raise filtering strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Operate and maintain a large, multi-vendor GPU fleet supporting demanding training and inference workloads to ensure reliability and performance.
Own on-call duties including incident response, runbook creation, and conducting postmortems to implement lasting fixes.
Develop and maintain internal tooling for fleet operation and collaborate with ML and platform teams to support large-scale AI workloads.
5+ years in infrastructure or site reliability engineering with at least 2 years operating large GPU clusters.
Proficiency in Python or Go for building and maintaining tooling.
Working knowledge across multiple infrastructure areas including storage, networking fabrics, and distributed training issues.
Demonstrated experience in on-call infrastructure ownership with successful incident management and follow-up.
Specialist with deep expertise in one of the five core focus areas while fluent enough in the others to triage and escalate appropriately.
Experienced in operating hybrid Kubernetes and Slurm environments for GPU workloads.
Familiarity with on-premise GPU deployment logistics such as power, cooling, and InfiniBand networking preferred but not mandatory.