





Niche SRE skillset and lesser-known employer reduces applicant density.
Requires specific SRE tooling and cloud experience, moderately limiting cross-industry portability.
Explicit 6–8 years plus mandatory GCP, Kubernetes, observability, and IaC tools makes screening strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own incident management including on-call, root cause analysis (RCA), and postmortems for critical production issues.
Define and monitor SLIs, SLOs, and error budgets to drive service availability and reduce mean time to recovery (MTTR).
Design and maintain observability frameworks and runbooks using tools like ELK, Prometheus, and Grafana, and improve automated system recovery and alerting.
6-8 years of Site Reliability Engineering or infrastructure engineering experience in cloud-native environments.
Strong hands-on expertise with GCP services (GKE, Load Balancing, VPN, IAM).
Proficient with observability tools: Prometheus, Grafana, ELK, Datadog; and containers/orchestration: Kubernetes, Docker.
Experience with incident management tools (PagerDuty, OpsGenie) and Infrastructure as Code (Terraform, Helm).
Experienced in leading cross-functional incident response including war-room coordination and postmortems.
Demonstrated capability in designing and implementing observability frameworks and telemetry adoption in production environments.
Strong operational focus on improving platform resilience, capacity planning, failover testing, and reliability reviews.