Senior Staff Site Reliability Engineer
NVIDIA CorporationMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessStrong employer brand and metro location increase competition despite seniority and specialized SRE requirements.
Requires niche AI GPU, inference, and database platform experience, limiting cross-industry transferability.
Explicit 8+ year requirement plus deep Kubernetes, control-plane, database, and GPU platform expertise narrows shortlist.
Job Description
Structured overview of role & requirementsAbout This Role
Define and lead the architecture and technical roadmap for a scalable enterprise AI runtime platform supporting deployment, operation, and scaling of AI applications, inference services, and databases across cloud and on-premises.
Design and build Kubernetes-based systems, including control-plane services, APIs, operators, automation for workload lifecycle management such as provisioning, configuration, upgrades, and recovery.
Lead technical initiatives, mentor engineers, and establish standards while improving performance, availability, and observability of large-scale AI inference services and database platforms.
Minimum Requirements
Bachelor's, Master's, or PhD in Computer Science, Engineering, or related fields, or equivalent experience.
8+ years of software engineering experience in distributed systems, cloud infrastructure, database platforms, or large-scale backend services.
Strong programming skills in Python, Go, C++, or Java with production-grade system delivery experience.
Proven experience designing scalable, highly available Kubernetes-based platforms and building control planes, Kubernetes operators, or workload lifecycle-management systems.
Ideal Candidate Profile
Experienced in leading technical strategy and cross-team influence for complex platform problems in large-scale AI runtime environments.
Deep knowledge of Kubernetes, AI inference workloads, GPU scheduling, relational and vector databases, observability, and cloud-native security.
Track record of building self-service platforms for application/infrastructure lifecycle management and contributions to open-source projects in Kubernetes, AI/ML infrastructure, distributed systems, or observability.
