Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand, metro location, and broad SRE skillset make standing out highly competitive.
Deep SRE skills plus regulated-industry experience demand industry-relevant background, limiting cross-industry transferability.
Mandatory 10+ years, regulated-environment experience, and strict SRE tech stack make shortlisting highly selective.
Job Description
Structured overview of role & requirementsAbout This Role
Build and maintain self-healing automation and author validated runbooks to enable automated or manual remediation of recurring failure modes.
Lead root-cause analysis post incidents, driving durable engineering fixes that reduce recurrence and improve system reliability.
Collaborate with Senior Principal SRE Engineer, automation engineering, and Operations to balance automation vs manual runbooks and validate operational outcomes.
Minimum Requirements
10+ years of progressive engineering experience with significant hands-on reliability engineering and self-healing automation in a multi-application production environment.
Experience in a regulated or audited environment (e.g., GxP, SOX, HIPAA, PCI) including change control and validated system compliance.
Hands-on proficiency with SRE toolset: observability (e.g., Splunk, Datadog), Terraform, Kubernetes, CI/CD pipeline hardening, and major public cloud platform (AWS, Azure, or GCP).
Bachelor's degree or higher in Computer Science, IT, or related engineering field.
Ideal Candidate Profile
Experienced individual contributor comfortable balancing engineering rigor with operational pragmatism, focusing on building durable, safe automated or manual remediation capabilities.
Strong track record of delivering measurable reliability improvements through automation, root cause fixing, and runbook management, especially in regulated or complex environments.
Comfortable working in collaboration with senior engineers and operations teams to enforce standards on automation graduation and system safety in a continuous operations environment.
