Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Define and measure service reliability with SLI, SLOs aligning to business and user expectations.
Enable rapid production deployment of new software/features balancing performance and risk per SLA.
Set up observability, alerting, incident management, documentation, runbooks, and automation including CI/CD pipelines for cloud and containerized systems.
Minimum Requirements
Experience with Kubernetes or similar container orchestration and monitoring tools such as Prometheus, Grafana, Open Telemetry Collector.
Proficiency in scripting/programming languages used for infrastructure automation (e.g., Terraform).
Strong knowledge of system architecture, networking, distributed systems, and disaster recovery for cloud systems.
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced in bridging development, cloud platform engineering, and product teams to implement SRE best practices.
Skilled in incident management including response, troubleshooting, and post-mortem analysis.
Familiar with deploying and managing scalable, fault-tolerant cloud-based systems with a focus on operational excellence and cultural advocacy of SRE principles.
