Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and ensure reliability of high-throughput, real-time ad-serving platform on GCP and Kubernetes with measurable SLOs and observability.
Build tools and automation to eliminate manual operational tasks and reduce toil, improving system health at scale.
Drive architectural decisions on reliability and scalability, lead incident response, capacity planning, and partner cross-functionally for resilient production readiness.
Minimum Requirements
Hands-on experience managing large-scale, cloud-native production systems using Google Cloud Platform and Kubernetes (GKE).
Proficiency in software development using Go, Python, Rust, or similar with experience shipping maintainable operational tools or services.
Experience with SLIs, SLOs, error budgets, incident management, infrastructure-as-code (Terraform), observability (Grafana, VictoriaLogs), and CI/CD/GitOps tooling (Argo CD, GitHub Actions).
Bachelor's degree in Computer Science or equivalent practical experience; Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced engineer with a strong background in cloud-native reliability engineering and production system scalability on GCP and Kubernetes.
Technical operator who efficiently blends software development skills with infrastructure management, focused on automation and measurable reliability.
A cross-functional collaborator who leads incident response, owns system health metrics, and influences platform architecture for high availability under critical workloads.
