





Tier-1 brand, metro location, mid-level SRE role with broad skills and on-call responsibility.
Core SRE skills (Dynatrace, Python, Kubernetes, Azure) are highly transferable across industries despite insurance preference.
Explicit 5+ years plus mandatory observability, programming, cloud, and on-call experience increases filter strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and operate automation and self-healing systems to reduce manual toil and improve production reliability for assigned services.
Own SLI/SLO instrumentation, error budget tracking, and Dynatrace observability including instrumentation, alert tuning, and dashboard design.
Act as a senior on-call escalation point and incident commander for critical production incidents, conducting root cause analysis and implementing durable fixes.
Bachelor's degree in Computer Science, Software Engineering, Information Technology, or equivalent.
Minimum 5 years of software engineering, platform reliability, or SRE experience with hands-on production service ownership.
Proven hands-on engineering skills in Python, Go, or equivalent for automation and self-healing workflows.
Experience with Dynatrace observability platform and participating in 24/7 on-call rotations as senior escalation for high-severity incidents.
Experienced senior-level SRE comfortable leading high-severity incident response and driving reliability improvements with measurable impact.
Expert in Dynatrace instrumentation, observability best practices, and automation-first reliability engineering.
Able to mentor junior engineers and contribute to the reliability roadmap while integrating AI-assisted tooling and chaos engineering into SRE workflows.