





Tier-1 brand, remote and popular SRE role with mid-level experience and broad skills, high competition.
Core cloud, Kubernetes, and SRE skills transfer across industries but require specific operational experience.
Explicit 5+ years plus Kubernetes, cloud, observability, SLOs, and scripting requirements force strict shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own monitoring, support, and reliability for large-scale GeForce NOW production services across cloud and datacenter environments.
Lead incident triage, troubleshooting, and resolution, including on-call rotation participation to ensure timely restoration of customer-facing services.
Design and develop automation, custom tools, and self-service solutions to improve operational efficiency and service reliability while collaborating with cross-functional teams.
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field (or equivalent experience).
5+ years experience supporting and operating mission-critical production services as SRE, Production Engineer, or similar role.
Strong understanding and operational experience with Kubernetes, containerization, microservices architecture, and distributed systems.
Experience with production incident management, automation scripting (Python, Go, Bash), observability platforms (Prometheus, Grafana, ELK/OpenSearch), and public cloud environments (AWS, Azure, GCP).
Experienced in supporting large-scale, customer-facing cloud or gaming services with strong Kubernetes troubleshooting and operational expertise.
Proficient in driving production incident response, postmortems, and operational excellence initiatives to improve long-term service reliability.
Advanced automation and programming skills focusing on reducing operational toil within complex, distributed system environments.