





Remote role, metro location, and broad required skillset increase applicant competition.
Core SRE skills transferable but heavy telecom networking and observability requirements raise domain specificity.
Explicit 8+ years and many mandatory technologies (GKE, Terraform, Grafana, Kafka) increase shortlist strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design and operation of global platform reliability including large-scale networking (Layers 1-7) and cloud-native observability on GCP.
Build and scale next-generation intelligent observability stack using Grafana Labs tools, ML-powered alerting, and automated incident response.
Provide technical leadership and mentorship across distributed infrastructure engineering, including Kubernetes (GKE), Kafka event streaming, and multi-database management.
8+ years of experience in Site Reliability Engineering, Production Engineering, or Distributed Systems infrastructure roles.
Deep expertise in full-stack networking (OSI Layers 1 through 7) including physical infrastructure, routing protocols, transport layer tuning, and advanced application protocols.
Advanced proficiency with Google Kubernetes Engine (GKE), Apache Kafka, PostgreSQL, AlloyDB, BigQuery, and Grafana Labs telemetry stack at scale.
Proficient in Terraform for multi-region GCP cloud infrastructure provisioning and Go/Python programming for custom tooling.
Demonstrated capability to work autonomously in fully remote, globally distributed engineering environments.
Strong strategic orientation in designing large-scale, cloud-native, AI-driven observability and reliability solutions on GCP.
Expertise in deep networking diagnostics combined with infrastructure-as-code and ML/AI for intelligent operations and automated remediation.