





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, Bengaluru metro location, and a visible senior DevOps title drive high applicant competition.
Requires specialized GPU and NVIDIA infrastructure knowledge, though core SRE skills remain transferable across cloud environments.
Explicit 8+ years, mandatory SRE/DevOps expertise and toolchain requirements yield high shortlisting strictness.
Lead operational readiness for NVIDIA Cloud Partner (NCP) infrastructure beyond initial deployment, focusing on Day 2 operations including health, observability, lifecycle management, and remediation of large-scale GPU-accelerated clusters.
Collaborate directly with NVIDIA Cloud Partners to establish and automate consistent operational procedures, monitoring, and validation for accelerated computing infrastructure in production.
Develop and implement frameworks, tooling, health metrics, SLOs, and automation to ensure reliable, scalable, and repeatable infrastructure operations supporting AI workloads.
Minimum 8+ years experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform or systems engineering supporting large-scale production environments.
Strong Linux-based distributed systems and cloud infrastructure operational experience, including Kubernetes and cluster lifecycle management.
Programming and automation experience with Python, Go, shell scripting or similar languages.
Education: BS, MS, or PhD in Computer Science, Computer/Electrical Engineering, or related technical field, or equivalent experience.
Experience managing large-scale GPU or accelerated computing infrastructure supporting AI training and inference workloads.
Proven ability to collaborate with strategic cloud partners, hyperscale providers, or managed AI cloud environments to operationalize infrastructure and maintain SLOs.
Deep expertise with NVIDIA technologies (e.g., DGX, CUDA, NVLink, GPU Operator) and infrastructure observability tooling (Prometheus, Grafana, OpenTelemetry), translating reference architectures into operational practices.