





Tier-1 employer, remote mid-level SRE role with broad cloud and Kubernetes requirements.
Requires deep SRE, cloud, and Kubernetes expertise, limiting easy industry transfer.
Explicit 5+ years, required SRE production experience, Kubernetes, cloud, and automation mandates.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own monitoring, support, and reliability maintenance for large-scale GeForce NOW production services across cloud and datacenter environments.
Lead incident triage, troubleshooting, resolution, and participate in on-call rotation to ensure customer-facing service uptime.
Develop and drive automation, custom tools, and observability improvements to enhance service SLOs, operational efficiency, and scalability.
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field (or equivalent experience).
5+ years experience supporting and operating mission-critical production services in live-site roles such as Site Reliability Engineer or Production Engineer.
Strong knowledge of Kubernetes, containerization, microservices, distributed systems, and public cloud platforms (AWS, Azure, GCP).
Proficiency in automation coding/scripting (Python, Go, Bash), familiarity with observability tools like Prometheus, Grafana, ELK/OpenSearch, and solid incident and operational process management.
Experienced with large-scale customer-facing cloud or gaming services with strong Kubernetes operational and troubleshooting skills.
Proficient in driving incident response, root cause analysis, postmortems, and continuous operational excellence.
Demonstrates strong automation development skills and a proactive approach to improving service reliability and scalability through engineering solutions.