





Metro location and broad platform skill requirements increase applicant density, but seniority and non-remote reduce competition.
Core SRE, cloud, and Kubernetes skills transfer easily across industries despite SaaS insurance context.
Explicit 8–12 year requirement and mandatory cloud, Kubernetes, and IaC skills make filters stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and operate highly reliable, scalable multi-tenant SaaS platform infrastructure with focus on automation and operational workflows.
Develop and maintain observability systems, define SLOs, lead incident management, and improve platform resilience and self-healing.
Collaborate with product teams to ensure production readiness, security compliance, and support 24x7 follow-the-sun operations.
8-12 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or related platform engineering role.
Strong programming skills in Python or Go; experience with AWS, Kubernetes (EKS), Docker, Helm, and Infrastructure as Code (Terraform/Terragrunt).
Experience with observability tools (Datadog, Prometheus, OpenTelemetry, CloudWatch) and incident management in microservices environments.
Knowledge of security and identity systems (SSO, SAML, OAuth), AWS IAM, and Kubernetes security is mandatory.
Experienced in designing and operating large-scale, multi-tenant SaaS platform environments with deep AWS and Kubernetes expertise.
Strong focus on automation and system reliability, actively reducing operational toil through tooling and self-healing solutions.
Able to mentor and influence across distributed engineering teams with a data-driven, systems-thinking approach.