Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead reliability outcomes for mission-critical network services focusing on availability, performance, and recoverability.
Conduct resiliency design reviews, lead major incident response, and own end-to-end problem management including root cause analysis and remediation.
Architect automation frameworks with Python, Shell, and Ansible; provide technical leadership in SD-WAN, SDA, routing, switching, firewalls, and promote AI-assisted reliability workflows.
Minimum Requirements
5+ years applied experience and formal training or certification in site reliability engineering concepts.
Extensive experience operating and engineering large-scale networks with strong troubleshooting skills.
Proven leadership in major incident coordination and root cause analysis with durable incident prevention.
Advanced automation skills in Python, Shell, and Ansible with demonstrated toil reduction and reliability improvement.
Ideal Candidate Profile
Experienced in managing SRE metrics including non-functional requirements and embedding SRE practices in design and delivery.
Demonstrated ability to independently drive delivery with strong ownership and stakeholder management.
Familiar with enterprise-authorized AI in improving SRE workflows and capable of evaluating AI-assisted operational recommendations for correctness and risk.
