Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Define technical vision and long-term architecture for reliability and infrastructure platform at scale.
Lead design and delivery of major platform initiatives: Kubernetes, observability frameworks, SLO programs, incident management.
Drive org-wide reliability standards, risk identification, remediation, and mentor senior SREs; influence technical hiring bar.
Minimum Requirements
Expert-level Kubernetes knowledge including internals, failure modes, and large-scale cluster management.
Proven 12+ years of hands-on experience with technical leadership on complex infrastructure initiatives at scale.
Expertise in SRE principles (SLIs, SLOs, error budgets) and observability stacks (OpenTelemetry, Prometheus, Grafana, Datadog).
Strong cloud engineering skills (AWS/GCP/Azure), distributed systems knowledge, incident management experience, and CI/CD expertise.
Ideal Candidate Profile
Experienced at Staff or Principal engineering level owning technical domains and setting engineering standards.
Capable of defining architectural direction and consensus-building without formal authority across cross-functional teams.
Demonstrates strategic focus on reliability as a product with measurable impact on organization-wide engineering operations.
