





Tier-1 employer, metro location, and a popular DevOps/SRE role drive high competition.
Requires specialized on-prem Kubernetes, SLURM, and HPC experience, limiting cross-industry transferability.
Explicit 7+ years requirement plus mandatory Kubernetes, on-prem, and Linux networking skills make shortlisting strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead and operate a 24/7 global Service Reliability Operations center focused on Production Kubernetes Services.
Develop and maintain automation, monitoring, alerts, and incident response to achieve near 100% service availability.
Perform large scale Kubernetes and Linux systems administration, drive root cause analysis and incident management for production cloud services.
7+ years administering large-scale production Kubernetes systems in Internet, Cloud, or Data Center environments, with strong preference for on-premises expertise.
Bachelor's degree in Computer Science, Engineering, Mathematics, or equivalent experience.
Advanced hands-on experience with Kubernetes, SLURM, large-scale cluster management, Linux system administration, networking (DNS, DHCP, IP Tables).
Experience working with CI/CD tools like Jenkins and ArgoCD, and scripting/programming in Python, Golang, or Rust preferred but not mandatory.
Experienced in architecting and scaling Kubernetes deployments for thousands of users in production-grade environments.
Strong systems troubleshooting skills with deep understanding of high-performance computing clusters including GPU/DPU hardware.
Able to lead complex incident management and collaborate cross-functionally to improve service reliability and automation.