





Remote mid-level SRE in a metro with broad, popular skill requirements increases candidate competition.
Core SRE, cloud, and observability skills are broadly transferable across industries.
Multiple mandatory skills, a 5+ year requirement, and specific cloud/container certifications increase shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Ensure high availability and reliability of products through building support systems, event management, and automation.
Own incident and root cause analysis processes to deliver permanent fixes and improve operational stability.
Collaborate across Product, Engineering, and Operations teams to optimize service readiness, deployment, and customer success initiatives.
Bachelor’s degree in Computer Science, IT, Engineering, or related field.
5+ years of experience in IT operations, Site Reliability Engineering, or infrastructure management, including incident management and root cause analysis.
Strong hands-on experience with Azure/Windows Environments, CI/CD pipelines, automation, monitoring, and infrastructure as code.
Relevant certifications recommended: VMware, AWS/Azure, MCSE, RHSA, CompTIA Security+, Certified Kubernetes Administrator (AKS or CKA).
Experienced in managing complex high-availability systems with strong expertise in cloud (Azure, AWS), containerization (AKS, Kubernetes), and automation tooling (Terraform, CI/CD).
Proficient in troubleshooting, problem management, and proactive monitoring using observability platforms such as Prometheus, Grafana, Dynatrace, Azure Monitor, and AWS CloudWatch.
Able to work cross-functionally with Product, Engineering, and Operations teams to enhance platform stability and performance at scale.