Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Manage and ensure stable, resilient, and reliable applications minimizing disruptions to Customer & Colleague Journeys (CCJ).
Define and enforce error budgets to balance risk and reliability while improving release processes for better velocity and system scalability.
Lead incident management and provide technical guidance to teams, proactively introducing innovations to meet service objectives.
Minimum Requirements
Experience with AWS tools (EMR, Airflow, S3, EKS, EC2, Lambda, CloudWatch) and DevOps tools (GitLab, GitLab CI, Artifactory).
Hands-on programming skills in Spark, Python, and shell scripting; experience with Docker, Prometheus, and Grafana.
Strong ability to debug complex data issues using Splunk, Spark, and CloudWatch audit logs; experience in incident and service management.
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced in site reliability engineering with a strong foundation in software engineering and reliability systems thinking.
Familiar with financial services domain and able to understand broader business impacts and risks related to services.
Proactive in stakeholder communication and skilled at leading incident responses and coaching teams in a mature DevOps environment.
