Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead resiliency design reviews and lead complex technical problem solving for medium to large products in site reliability engineering.
Champion site reliability culture and lead initiatives to improve application and platform reliability with data-driven analytics.
Serve as main contact during major incidents and guide use of enterprise AI tools to enhance incident triage, troubleshooting, and reliability workflows.
Minimum Requirements
5+ years of applied experience in site reliability engineering with formal training or certification in SRE concepts.
Proficiency in at least one programming language (e.g., Java, Python, Java Spring Boot).
Experience with observability tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk), CI/CD tools (e.g., Jenkins, GitLab, Terraform), and container orchestration (e.g., ECS, Kubernetes, Docker).
Work Experience Required: 5+ years applied site reliability engineering experience.
Ideal Candidate Profile
Technical leadership ability to manage cross-functional teams and break down complex engineering problems in a large-scale banking technology environment.
Strong expertise in reliability engineering best practices, SDLC, secure development, DevOps/CI/CD, and AI-assisted operational workflows.
Experience evaluating and validating AI-driven operational recommendations while managing risk and security in enterprise systems.
