Senior Site Reliability Engineer, Production Engineering
NVIDIA CorporationMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand, remote role, metro location, and broad SRE demand create high competition.
On-prem Kubernetes, bare-metal infrastructure and HPC/GPU specialization reduce transferability across industries.
Explicit 7+ years plus mandatory large-scale Kubernetes, bare-metal, and HPC experience makes filters highly strict.
Job Description
Structured overview of role & requirementsAbout This Role
Support and automate production Kubernetes services within a 24/7 global service reliability operations center, including split-weekend shifts.
Perform large-scale Kubernetes administration and systems security monitoring to maintain service SLAs and reliability.
Lead incident management through proactive monitoring, root cause analysis, and coordination with SMEs to minimize incident frequency and duration.
Minimum Requirements
7+ years experience administering large-scale production Kubernetes systems, preferably with on-premises expertise.
Bachelor's degree in Computer Science, Engineering, Mathematics, or equivalent experience.
Strong Linux system administration skills including DNS, DHCP, IP Tables, routing, and firewalls on large-scale bare-metal infrastructure.
Experienced with Kubernetes, SLURM, CI/CD tools (Jenkins, ArgoCD), and familiar with GPU/DPU hardware in high-performance computing clusters.
Ideal Candidate Profile
Experienced in architecting and deploying large-scale Kubernetes environments supporting thousands of users.
Technically proficient with deep knowledge of Kubernetes, large-scale cluster management, and system security in internet or cloud data center contexts.
Comfortable driving incident response and cross-functional communication within fast-paced and complex production engineering teams.
