Senior Site Reliability Engineer, Production Engineering
NVIDIA CorporationMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand and Bangalore metro increase candidate density, but niche GPU/Kubernetes on-prem skills limit applicants.
Highly specialized on-prem Kubernetes and HPC/GPU cluster skills reduce cross-industry transferability.
Mandatory 7+ years Kubernetes/SRE experience and on-prem GPU/bare-metal expertise enforces strict screening.
Job Description
Structured overview of role & requirementsAbout This Role
Lead global Service Reliability Operations center focused on Cloud product support, ensuring near 100% service availability.
Manage 24/7 Production Kubernetes Services including large-scale Kubernetes and systems administration, automation, and security monitoring to maintain SLAs.
Proactively monitor and respond to incidents using alerts and observability tools; lead root cause analysis and incident management calls for timely resolution.
Minimum Requirements
7+ years experience administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center environments (strong preference for on-prem).
BS degree in Computer Science, Engineering, Mathematics, or equivalent experience.
Advanced hands-on skills in Kubernetes, SLURM, cluster management; strong Linux system administration including DNS, DHCP, core networking (IP Tables, routing, firewalls).
Experience with CI/CD tools like Jenkins or ArgoCD; scripting/programming knowledge in Python, Golang, or Rust preferred but not mandatory.
Ideal Candidate Profile
Experienced in architecting, building, and deploying Kubernetes for large-scale environments serving thousands of users.
Deep technical expertise with on-prem bare-metal infrastructure and high-performance computing clusters, including GPU/DPU hardware familiarity.
Comfortable leading incident management and cross-functional communications in a 24/7 global support environment with a focus on automation and rapid issue resolution.
