





Tier-1 brand, remote role, mid-level SRE, and metro location increase applicant density.
Technical SRE skills are moderately transferable, but require specific operational and cloud experience.
Mandatory 5+ years SRE experience, Kubernetes, cloud, observability, and automation create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own monitoring, operational support, and reliability of large-scale GeForce NOW cloud gaming production services across cloud and datacenter environments.
Lead incident management including triage, troubleshooting, root cause analysis, and blameless postmortems to improve service reliability and operational workflows.
Design and develop automation tools and self-service solutions to enhance productivity, observability, and scalability of Kubernetes-based services.
Bachelor’s degree in Computer Science, Computer Engineering, IT, or related field (or equivalent experience).
5+ years experience in site reliability engineering, production engineering, or similar role supporting mission-critical production services.
Strong knowledge of containerization, microservices, Kubernetes, distributed systems, and cloud platforms (AWS, Azure, GCP, or equivalent).
Proficiency in scripting or programming languages like Python, Go, or Bash and experience with observability tools such as Prometheus, Grafana, and ELK/OpenSearch.
Experienced in managing large-scale, customer-facing cloud or gaming services with strong operational and troubleshooting skills in Kubernetes environments.
Practiced in applying SRE principles including SLOs, SLIs, error budgets, incident response, and operational excellence initiatives.
Hands-on engineer capable of building automation and monitoring tooling to improve service reliability and engineer productivity.