





Remote mid-level SRE role with broad requirements and popular title increases applicant competition.
SRE skills are broadly transferable across industries, so candidate backgrounds are easily adaptable.
Explicit 5+ years and mandatory cloud, Kubernetes, observability, and coding skills make filters stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead incident response as Incident Commander, coordinating cross-functional teams during major incidents and managing on-call rotations as the primary technical lead.
Define, measure, and defend Service Level Objectives (SLOs) and Error Budgets by instrumenting high-cardinality metrics and distributed tracing for distributed systems.
Develop production-quality code (Java, Go, Python) to build automation and self-healing mechanisms, while partnering with product teams to embed reliability, scalability, and observability best practices from design phase.
Minimum 5 years of experience in Site Reliability Engineering or Backend Engineering with production coding expertise in Java, Go, Rust, or Python.
Strong knowledge of distributed systems architecture, microservices fundamentals, event-driven architectures, and scaling principles.
Experience with Google Cloud Platform or equivalent cloud providers, Kubernetes production workloads (GKE/EKS), and cluster/infrastructure troubleshooting.
Experience with observability tools (OpenTelemetry, Prometheus, New Relic, Datadog, SigNoz) and production data stores (PostgreSQL, MySQL) and streaming platforms (Kafka, RabbitMQ).
Experienced engineer with proven ownership of complex distributed systems reliability and incident management at scale.
Skilled coder comfortable building internal tools and automation to reduce manual interventions in production environments.
Strong collaborator able to work closely with product engineering during design to implement reliability and observability patterns from day one.