Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)
SkyflowMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Build and maintain automated, scalable cloud infrastructure provisioning and management across multi-tenant and BYOC deployments using Go and Python.
Own platform reliability including Kubernetes clusters, service mesh, databases, and messaging systems with operational responsibilities: on-call, incident response, root-cause analysis.
Develop Infrastructure-as-Code and tooling (Terraform/OpenTofu, Helm, GitOps) to enable self-service platform capabilities for engineering teams.
Minimum Requirements
4+ years in platform engineering, infrastructure engineering, DevOps, or SRE with production cloud infrastructure ownership.
Strong programming skills in Go and/or Python focused on tested, maintainable infrastructure automation.
Hands-on Kubernetes experience in production, including workload scheduling, networking, autoscaling, and upgrades.
Experience with at least one major public cloud provider (AWS, GCP, or Azure).
Ideal Candidate Profile
Experience operating complex B2B cloud platforms with compliance, security, and uptime SLAs, including dedicated/BYOC deployments.
Proven ability to build internal developer tools (CLIs, provisioning frameworks, automation pipelines) to reduce infrastructure toil.
Strong incident management skills with capacity to debug distributed systems under pressure and lead root-cause analysis with corrective actions.
