





Strong Tier-1 brand, mid-level generalist SRE, metro location, broad cloud/Kubernetes skills make hiring highly competitive.
SRE and cloud-operational skills are transferable across industries but require domain-specific tooling experience.
Explicit 5+ years requirement plus mandatory Kubernetes, cloud, SRE practices, and automation skills imply strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own reliability, availability, and performance of large-scale GeForce NOW cloud gaming production services.
Lead incident response, triage, troubleshooting, root cause analysis, and drive corrective actions for complex infrastructure and application issues.
Develop automation, custom tools, and improvements to operational processes to enhance service reliability and efficiency, including Kubernetes-based services.
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related field (or equivalent experience).
5+ years supporting and operating mission-critical production services as Site Reliability Engineer (SRE) or similar role.
Strong expertise with containerization, microservices, Kubernetes, and distributed systems in production environments.
Hands-on programming/scripting skills in Python, Go, Bash or similar for automation and experience with public cloud platforms (AWS, Azure, GCP).
Experienced in large-scale customer-facing cloud or gaming service operations with strong SRE practices.
Deep operational knowledge of Kubernetes ecosystem and modern observability platforms (Prometheus, Grafana, ELK/OpenSearch).
Proven ability to drive incident management, postmortem processes, and promote operational excellence initiatives.