Site Reliability Engineer, AVP
Deutsche BankMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Define and improve SLIs, SLOs, alerting standards, and error budgets for an on-premises Kubernetes platform.
Investigate and resolve complex production issues across Kubernetes, Linux, networking, storage, and platform dependencies with ownership of incident response and mitigation.
Automate operational tasks to reduce toil, enhance platform reliability, capacity planning, disaster recovery, and work collaboratively with global engineering teams in a 24x7 support model.
Minimum Requirements
Bachelor’s degree in a technical or engineering discipline.
10–12 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps, or closely related infrastructure role.
Strong Kubernetes expertise including cluster operations, upgrades, troubleshooting, networking, storage, security, Helm, Operators in bare-metal or private-cloud.
Work Experience Required: 10–12 years in relevant infrastructure roles (SRE, Production Engineering, DevOps).
Ideal Candidate Profile
Demonstrated ability to manage SLIs/SLOs, incident reduction, capacity planning, performance optimization, and operational excellence with an SRE mindset.
Experienced with Kubernetes platforms, observability tools (Prometheus, Grafana, Splunk), automation scripting (Python, Bash, Ansible), and CI/CD/GitOps tools (Argo CD, Jenkins).
Comfortable working independently on complex problems and collaborating globally with clear communication, including experience in a 24x7 on-call follow-the-sun support environment.
