





Popular mid-level SRE role, metro context, broad skillset and 2+ experience increases competition.
Core SRE and cloud skills are highly transferable across industries.
Explicit 2+ years requirement plus specific monitoring, cloud, and SRE tooling raises filtering strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Ensure availability, reliability, and performance of mission-critical applications and services within SLA/SLO targets.
Lead monitoring, incident response, problem management, and automation initiatives across hybrid cloud and enterprise environments.
Collaborate with multiple technical teams and business stakeholders to improve operational readiness, reduce incidents, and enhance system resiliency.
Minimum 2 years experience in Site Reliability Engineering, Production Support, Systems Engineering, DevOps, NOC, or Command Center Operations.
Experience with monitoring tools like AppDynamics, Dynatrace, Datadog, Splunk, Azure Monitor, and ServiceNow Event Management.
Strong understanding of Incident, Problem, Change, and Service Level Management.
Knowledge of major cloud platforms: Microsoft Azure, Google Cloud Platform (GCP), and AWS.
Experienced in 24x7 operational support and major incident management in a Global Command Center or similar environment.
Skilled with scripting and automation using PowerShell, Python, Bash, or APIs to improve operational efficiency and implement self-healing.
Familiar with SRE concepts such as SLI/SLO/SLA management, error budgets, and resiliency engineering, preferably in retail, hospitality, payments, or enterprise SaaS domains.