





Senior, niche SRE role in a metro reduces applicant density despite moderate visibility.
SRE requires specific production, SLO, and on-call experience, limiting cross-industry transferability.
Explicit 8+ years and mandatory SRE ownership with SLOs and incident command make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own production reliability and uptime for a high-speed real-time platform combining voice, desktop, intelligence, and AI, with zero tolerance for retries during agent calls.
Be accountable for Service Level Objectives (SLOs), error budgets per tenant/service, incident response, blameless postmortems, production scaling, capacity planning, and observability metrics tracking (p50/p95/p99 per event hop).
Participate in on-call rotations and collaborate closely with Platform Foundation teams on observability and tenancy isolation, operating in a fast-paced startup environment with weekly deployments.
Minimum 8+ years experience operating production systems at scale with ownership of SLOs, error budgets, and incident command roles.
Experience in incident response, production scaling, capacity planning, and observability of high-throughput, real-time systems.
Ability to work with rapid deployment cycles and agile sprint cadences (weekly deploys, 1-week sprints).
Work Experience Required: Minimum 8+ years in relevant production system reliability roles.
Experienced in building and maintaining SLOs and error budgets for multi-tenant real-time platforms with a high uptime requirement.
Operates effectively in fast-paced, agile startup environments with frequent releases and evolving tech stacks.
Demonstrates calm and control during incidents, data-driven with dashboards and drills, and embraces new technologies and AI tools to enhance reliability and velocity.