Sr Lead SRE - Reliability Engineering & Problem Management
JPMorgan Chase & Co.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Serve as technical authority for root cause analysis (RCA) reviews and post-incident investigations, ensuring corrective actions address systemic causes.
Lead reliability maturity assessments, define and govern SLOs/SLIs, and drive adoption of AI-assisted reliability workflows across incident lifecycle and SDLC/toolchain practices.
Evaluate detection gaps, resilience weaknesses, and architecture; recommend resilience patterns and monitor corrective action effectiveness to reduce incident recurrence.
Minimum Requirements
5+ years applied experience in Infrastructure Engineering, Site Reliability Engineering, Production Engineering, Systems Engineering, or Software Engineering with hands-on experience leading/supporting critical incident investigations.
Formal training or certification on security engineering concepts.
Advanced expertise in Network Engineering, Cloud Infrastructure, Linux/Windows Platforms, Middleware, Databases, Storage, Application Architecture, DevOps Toolchains, Distributed Systems, and Enterprise Monitoring platforms (e.g., Splunk, Dynatrace, Grafana).
Strong proficiency in SLO/SLI engineering, reliability metrics, distributed tracing, telemetry, error budget management, AIOps platforms, and experience with Root Cause Analysis methodologies (Five Whys, Fault Tree Analysis, etc.).
Ideal Candidate Profile
Experienced senior engineer capable of challenging senior stakeholders constructively and influencing without direct authority, translating technical findings into executive communications.
Deep practical experience in large-scale, highly regulated, mission-critical environments applying systemic and evidence-based problem management.
Demonstrated ability to integrate and govern AI-enabled workflows for reliability engineering, ensuring auditability, security, and data sensitivity compliance.
