





Strong Tier-1 brand, mid-level SRE title, metro hiring, and broad skill requirements.
SRE expertise in cloud, Kubernetes, and observability is highly transferable across industries.
Explicit 5+ years requirement and mandatory Kubernetes, cloud, observability, and automation skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own monitoring, support, and maintenance of large-scale GeForce NOW production services across cloud and datacenter environments to ensure reliability, availability, and performance.
Drive tools and service development including automation and custom solutions to improve GeForce NOW service SLOs and operational efficiency.
Lead incident response, root cause analysis, and collaborate with multiple teams to improve operational readiness and service resilience.
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related field (or equivalent experience).
Minimum 5 years of experience operating mission-critical production services in Site Reliability Engineering or similar roles.
Hands-on experience with Kubernetes, containerization, microservices architecture, and ability to troubleshoot complex production issues.
Proficiency in programming/scripting languages like Python, Go, or Bash; experience with observability tools (Prometheus, Grafana, ELK/OpenSearch); and experience operating services in public cloud environments (AWS, Azure, GCP).
Experienced with large-scale customer-facing cloud or gaming service operations with a strong focus on production incident response and operational excellence.
Advanced Kubernetes operational and troubleshooting skills, including knowledge of Kubernetes ecosystem components and best practices.
Demonstrated ability to lead improvements in service reliability using SLOs, SLIs, error budgets, and effective automation to reduce operational toil.