





Remote role, mid-level SRE title, broad cloud/Kubernetes/observability requirements increase competition.
SRE skillset is broadly transferable across tech companies but requires specific cloud and infra experience.
Explicit 5-7 years plus mandatory GCP, Kubernetes, observability, databases, and coding requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead incident response as Incident Commander during major severity incidents, coordinating cross-functional teams and managing war rooms.
Develop and maintain automation and self-healing solutions by writing production-ready code in Java, Go, Python, or Rust to reduce manual intervention.
Define, measure, and defend Service Level Objectives (SLOs) and Error Budgets; partner with product engineering teams to incorporate reliability, scalability, and observability into new services.
5-7 years of experience in Site Reliability Engineering or Backend Engineering with strong coding skills in Java, Go, Rust, or Python.
Deep understanding of distributed systems architecture, microservices, and event-driven architectures.
Experience with Google Cloud Platform (GCP) or equivalent cloud providers and production Kubernetes operations (GKE/EKS).
Proficiency in observability tools (OpenTelemetry, Prometheus, New Relic, Datadog, or SigNoz) and experience with production databases (PostgreSQL, MySQL) and streaming platforms (Kafka, RabbitMQ).
Experienced SRE capable of balancing immediate operational stability with long-term system scalability in a hyper-growth environment.
Skilled in proactively designing and enforcing reliability and observability patterns at service design stage.
Comfortable operating in a complex distributed system environment with strong cross-team coordination and incident leadership skills.