





Tier-1 employer, mid-level SRE role, and metro location create high applicant competition.
Core SRE skills are transferable, though AIOps and OpenShift bias toward platform-focused employers.
Explicit 5+ years plus mandatory SRE, observability, IaC, cloud, and AIOps skills indicate high filtering strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead the definition and enforcement of SLOs and Error Budgets for critical services to balance feature velocity and system stability.
Drive adoption of AIOps solutions including anomaly detection and predictive alerting to reduce incident volume and improve MTTR.
Design and implement automation solutions to eliminate manual toil, build intelligent workflows with LLM integrations, and mentor junior engineers.
5+ years of experience in Site Reliability Engineering or similar role.
Proficiency in Python or Go programming and automation using Ansible.
Expert knowledge of observability tools such as TIGK Stack and Prometheus.
Deep understanding of Linux, Kubernetes/OpenShift, and public cloud infrastructure (AWS/Azure/GCP).
Technical leader capable of defining and implementing SLOs, Error Budgets, and incident management lifecycles.
Experience with advanced automation including Model Context Protocol and LLM-based operational workflows.
Strategic thinker able to translate business goals into technical reliability roadmaps and mentor others on complex systems engineering.