





Hybrid/remote, metro location and senior SRE demand create moderate candidate competition.
Specialized SRE, incident management, and cloud/Kubernetes skills limit cross-industry transferability.
Explicit 8+ years, mandatory SRE, Kubernetes, cloud, languages and leadership make screening stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead critical incident response as Incident Commander ensuring end-to-end restoration during complex production outages.
Drive engineering best practices to improve platform reliability, scalability, and customer experience, including automation and self-healing systems.
Mentor senior engineers and influence cross-functional strategy and culture around site reliability and operational excellence.
8–12+ years in Site Reliability Engineering or Production Operations.
Bachelor's degree or equivalent.
Expertise in Linux, Kubernetes, Cloud technologies, Networking, Databases, and programming with Python, Go, Bash, or Java.
Experience leading major incident management and mentoring engineering teams.
Experienced leader with a track record of reducing critical incidents and improving enterprise-wide reliability in large-scale SaaS environments.
Proven ability to drive automation and operational excellence across cross-functional teams including Engineering, DevOps, and Support.
Comfortable managing incident response under pressure and influencing senior leadership and engineering priorities globally.