





Broad observability and cloud skillset requirement increases applicant competition and relevance.
SRE and observability skills transfer across industries but require cloud and tooling experience, so moderately sensitive.
Many mandatory platform and tooling skills (Kubernetes, Azure/GCP, Prometheus, Terraform, Go/Python) drive strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and operate a unified observability platform for end-to-end visibility across Restaurants, Retail, and Payments environments.
Design and implement enterprise observability solutions on Azure, GCP, Kubernetes, and hybrid environments, including dashboards and telemetry standards.
Define and implement SLIs, SLOs, error budgets, improve alerting and incident response, and integrate observability with automation and operational workflows.
Experience in Site Reliability Engineering and managing Kubernetes (AKS/GKE).
Proficient with Azure and Google Cloud Platform.
Skills with observability tools such as Grafana, Dynatrace, Prometheus, OpenTelemetry or similar.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in designing and operating large-scale, multi-cloud observability platforms with focus on reliability and automation.
Skilled in implementing telemetry, SLIs/SLOs, error budgets, and operational analytics for proactive service management.
Capable of collaborating across Product Engineering, Infrastructure, Security, and Operations teams for improving platform reliability and readiness.