





Strong Tier-1 brand and metro location, but senior niche SRE skills limit applicant density.
Requires deep SRE, observability, hybrid cloud, and distributed systems expertise, limiting cross-industry transferability.
Explicit 15+ years, 7+ years leadership, and deep SRE, observability, IaC requirements enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design and evolution of scalable, reliable distributed systems applying SRE principles across infrastructure and application layers.
Drive automation to reduce manual effort, integrate reliability into software delivery via CI/CD, and establish operational metrics for service health.
Lead incident management including root cause analysis and implement observability solutions (metrics, logs, traces) to improve detection, reliability, and system performance.
Bachelor’s degree in Computer Science, Engineering, or related discipline (or equivalent experience).
15+ years in systems engineering with emphasis on site reliability, large-scale system operations, and software engineering in complex enterprise/cloud environments.
7+ years in technical leadership roles managing cross-functional initiatives and delivering complex projects.
Proficiency in programming languages like Python, Go, Java, Ruby; experience with hybrid environments, containerization, observability tools, and Infrastructure as Code with CI/CD pipelines.
Experienced in managing highly reliable, scalable distributed systems with strong operational and automation focus within complex enterprise or cloud contexts.
A technical leader skilled in cross-team collaboration, capable of translating reliability strategies and metrics into actionable insights for diverse stakeholder groups.
Practitioner of SRE best practices including incident management, post-incident reviews, capacity planning, chaos engineering, and continuous reliability improvement.