Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure high availability, performance, and recoverability of JumpCloud's critical systems and APIs across AWS and GCP.
Design and implement automation, observability frameworks, and reliability standards to reduce operational toil and improve incident detection and response.
Manage production Kubernetes clusters with GitOps workflows, provision multi-cloud infrastructure using Terraform, and develop disaster recovery dashboards and automation following RTO and RPO targets.
Minimum Requirements
5+ years professional software engineering experience in SRE, DevOps, or Platform Engineering with 24/7 mission-critical systems.
Proficiency in Python and/or Go for developing SRE tools and automation.
Production experience managing Kubernetes (EKS), GitOps tools like Argo CD, and Infrastructure as Code using Terraform.
Experience operating cloud workloads on AWS (including EKS, IAM, VPC networking) or GCP and practical knowledge of monitoring, incident management (e.g., Datadog, PagerDuty), and disaster recovery.
Ideal Candidate Profile
Experienced in building scalable cloud-native infrastructure and defining reliability standards within fast-paced, mission-critical environments.
Strong operational expertise in Kubernetes ecosystems and multi-cloud infrastructure management using automation and Infrastructure as Code.
Demonstrated ability to improve system resilience and operational efficiency through coding, incident management, and cost optimization practices (FinOps).
