





Strong employer brand, metro location, broad SRE skillset and generalist senior title increase candidate competition.
Core SRE cloud and Kubernetes skills are widely transferable across industries despite healthcare context.
Mandatory on-call SRE experience, cloud, Kubernetes and IaC requirements make screening technically stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and maintain scalable, fault-tolerant architectures in AWS and Azure, focusing on reliability, automation, and operational excellence for healthcare platforms.
Lead incident management: detect anomalies, perform root cause analysis, improve MTTR, and participate in a 24x7 on-call rotation to ensure system uptime and reliability.
Drive continuous improvement by defining SLIs/SLOs, optimizing capacity and costs, enhancing monitoring/alerting systems, and mentoring engineers on SRE best practices.
Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
Experience in site reliability engineering, software engineering, or related roles including production on-call experience.
Proficiency with AWS and/or Azure cloud platforms, including Kubernetes environments (EKS, AKS, GKE).
Proficiency in scripting languages for automation (e.g., Python) and experience with observability and incident management tools.
Strategic operator skilled at embedding reliability early in the software development lifecycle across cross-functional teams.
Experienced troubleshooting and system resilience expert in distributed cloud environments with a focus on fault tolerance and automation.
Senior-level SRE with ability to lead incident response, capacity planning, and mentor mid-level engineers to elevate organizational SRE maturity.