





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Strong brand and metro location but senior, specialized SRE skills reduce broad applicant density.
Role requires specialized SRE, cloud, and AI-ops experience, making cross-industry transferability limited.
Explicit 10+ years, mandatory SRE/incident leadership and specific cloud/tooling requirements create strict filters.
Lead end-to-end major incident management including triage, coordination, decision making, and executive communication across global zones.
Set technical direction and drive adoption of SRE initiatives to improve reliability, scalability, and developer efficiency across enterprise systems.
Design, build, and operate distributed and cloud-native systems, develop automation for incident detection and remediation, and apply AI techniques to enhance operational workflows.
10+ years experience in Site Reliability Engineering, Production or Platform Engineering, or Incident Management with technical leadership at scale.
BS or MS in Computer Science, Engineering, or related technical field, or equivalent experience.
Proven experience as Incident Commander in high-availability environments and deep knowledge of distributed systems reliability principles (SLIs/SLOs, error budgets).
Proficiency in one programming language (Python, Go, Java), public cloud platforms (AWS/Azure/GCP), Kubernetes/Docker, infrastructure-as-code tools, observability tools (Prometheus, Grafana), and networking fundamentals.
Experienced in balancing real-time incident leadership with scalable reliability solutions across distributed and multi-region teams.
Demonstrated ability to build and apply AI/ML tools for incident management and automation that measurably reduce detection and recovery times.
Track record of ownership and community engagement including contributions to open-source infrastructure or participation in SRE ecosystems.