





Mid-level SRE in Bangalore with popular SRE title but niche GPU/HPC specialization reduces applicant density.
Highly specialized GPU/HPC SRE skills limit transferability across industries.
Explicit 5+ years, 2+ years GPU-cluster requirement, on-call and language/tool mandates create strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Operate and maintain a large, multi-vendor GPU fleet supporting both long-running training jobs and low-latency inference services with a focus on reliability.
Own end-to-end GPU fleet provisioning, observability, capacity management, and health, including meaningful on-call rotation and incident management with durable postmortems.
Develop and maintain internal tooling, collaborate closely with ML and platform teams to ensure smooth operation of training and serving workloads.
5+ years experience in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.
Proficiency in Python or Go for building and maintaining internal tools.
Working fluency across multiple specialized GPU infrastructure focus areas (storage, fabric, network, firmware, kernel, training failures) to triage and route issues correctly.
Demonstrated on-call ownership of critical infrastructure with a track record of effective postmortem-driven fixes.
Specialist depth in one of the key infrastructure domains (storage, fabric, driver/firmware, kernel, distributed training), with working knowledge of others to handle complex incident triage.
Experienced operating and troubleshooting heterogeneous, multi-vendor GPU clusters with specialized reliability challenges, beyond basic Kubernetes administration.
Operationally mature with a proven ability to write durable runbooks, lead postmortems, and partner effectively across ML, platform, and datacenter teams.