Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Operate and maintain large-scale AWS and/or GCP cloud infrastructure with a focus on deployment, incident response, and service reliability.
Develop automation solutions using Go, Python, Terraform, and related tools to reduce operational toil and enable safe, repeatable infrastructure deployments.
Participate in on-call rotations, monitor SLIs/SLOs, handle incident responses, and collaborate with engineering teams to migrate workloads into modern platform capabilities.
Minimum Requirements
Hands-on experience operating production services in AWS and/or GCP environments.
Proficiency in software development with Go and/or Python plus experience using Terraform and Helm for infrastructure as code.
Solid, practical understanding of Kubernetes and container runtime troubleshooting.
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Has a strong operational and automation mindset with preference for ‘if you have to do it more than once, automate it’.
Experienced in managing customer-facing production systems with familiarity in SRE principles such as SLOs and blameless post-mortems.
Comfortable working in a hybrid environment and collaborating with globally distributed teams on infrastructure and platform engineering initiatives.
