





Tier-1 brand and broad SRE/cloud/Kubernetes requirements create moderate applicant competition.
Core SRE and platform skills are transferable, though utility/regulatory experience increases industry specificity.
Explicit 14+ years and mandatory SRE/cloud/Kubernetes experience create highly stringent shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the production stability and change management for global GridOS SaaS products, including final approval authority for all deployments.
Lead design and standardization of cloud infrastructure provisioning and software delivery platforms across multi-region architecture with 24/7 global operational coverage.
Serve as Lead Incident Commander for Sev1/Sev2 events, manage business continuity plans, and lead FinOps for cloud cost optimization and capacity planning.
14+ years overall professional experience; 8-10+ years in SRE, Platform Engineering, or Production Support for large-scale distributed SaaS.
Deep expertise in AWS core services (EC2, EKS, RDS, S3, IAM), Kubernetes/EKS multi-region architectures, ArgoCD, GitHub Actions, Terraform, Ansible, and observability tools (Prometheus, Grafana, Splunk/Datadog).
Master's degree in STEM OR Bachelor's in STEM with over 10 years relevant experience.
Experience managing/mentoring global SRE teams and influencing senior leadership stakeholders.
Experienced senior leader in SRE or Platform Engineering with proven ability to drive organizational change and enforce operational discipline at VP stakeholder level.
Skilled in designing and operating global SaaS infrastructure with multi-region cloud deployments and high reliability targets.
Hands-on operator with expertise in incident command, blameless post-incident analysis, and proactive system reliability engineering.