Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own the reliability and health of cloud applications through designing observability dashboards, alerting based on SLIs/SLOs, and resolving production incidents including root cause analysis.
Operate and improve Kubernetes-based compute platform and cloud infrastructure (Azure/AWS) for scalable, reliable systems.
Drive reliability best practices, automation of manual operational work, and support safe, quick shipping via CI/CD pipelines.
Minimum Requirements
Strong hands-on Kubernetes expertise is mandatory.
9+ years of relevant hands-on Site Reliability Engineering experience.
Practical experience with SRE principles including SLIs, SLOs, and error budgets on real systems.
Deep experience with cloud networking fundamentals and observability tools (e.g., OpenTelemetry, Prometheus, Grafana, Datadog).
Ideal Candidate Profile
Experienced in handling live production incidents with troubleshooting and improving mean time to resolution (MTTR).
Strong background operating distributed systems at scale including knowledge of failure modes and remediation.
Familiar with AI-assisted engineering tools for automating investigations and root cause analysis in production environments.
