Staff Network Reliability Engineer- Cloud Operations
SkyloMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and ensure 24x7 health and reliability of Skylo’s hybrid cloud production infrastructure including GKE clusters, on-prem Kubernetes, storage, and database systems.
Serve as L3 escalation authority diagnosing and resolving complex cloud infrastructure failures, owning runbooks, incident RCA, and error budget management tied to network SLA commitments.
Drive SLO definition, toil reduction automation, capacity planning, and mentor senior engineers while interfacing with platform engineering and network implementation teams for GitOps and infrastructure changes.
Minimum Requirements
8–10+ years in infrastructure engineering, site reliability, or cloud operations supporting Kubernetes-at-scale environments with direct on-call ownership.
Expertise in multi-cluster Kubernetes (GKE/EKS), hybrid cloud operations (public cloud and bare-metal/private cloud), and production observability tooling (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry).
Strong skills in database reliability (PostgreSQL streaming replication, Redis cluster operations), GitOps tooling (ArgoCD/Flux, Helm, Terraform/Ansible), and SRE fundamentals including SLO/SLA management and incident on-call rotations.
Work Experience Required: 8–10+ years; Notice Period: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced senior-level engineer comfortable operating and troubleshooting complex hybrid cloud telecom infrastructure with ownership of mission-critical production networks.
Proven track record in SRE practices focused on reliability engineering, error budget management, automation, and capacity forecasting in multi-cloud and private-cloud environments.
Strong cross-functional collaborator with the ability to define operational standards, runbooks, and infrastructure requirements influencing engineering roadmaps and effective escalation management.
