





Strong employer brand, metro location, and broad multi-discipline requirements increase applicant competition moderately.
Specialized HA/DR, cloud and networking skills transferable across SaaS and enterprise environments.
Mandatory 8+ years plus specific HA, IaC, Kubernetes, and networking expertise make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own architectural strategy and technical execution for platform high availability, resiliency, and disaster recovery lifecycle across cloud and on-prem environments.
Design, build, and deliver solutions including active-active clustering, global load balancing, and real-time database replication to eliminate single points of failure.
Lead cross-functional coordination for multi-region disaster preparedness and establish chaos engineering practices with automated fault-injection and failover drills.
8+ years of High Availability engineering experience across cloud (AWS, GCP, or Azure) and on-premises environments.
Hands-on expertise with real-time data replication, database clustering, and automated zero-downtime traffic failover.
Advanced proficiency with Infrastructure as Code tools like Terraform or Ansible.
Deep knowledge of Linux internals, Kubernetes container orchestration, and hybrid network topologies including BGP routing, DNS management, Anycast, and CDNs.
Demonstrated ability to translate uptime requirements into executable technical roadmap and personally execute the engineering work.
Proven experience influencing cross-functional teams to embed resiliency, self-healing, and automated recovery protocols.
Experienced in managing large-scale distributed systems with emphasis on fault tolerance, multi-region disaster recovery, and cloud/on-prem hybrid infrastructure.