





Metro location and recognizable employer increase competition, but niche SRE tooling narrows qualified applicant pool.
Core SRE skills like IaC, monitoring, and CI/CD are broadly transferable across industries.
Mandatory SRE tooling and platform experience imply strict technical filters and vetting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and improve site reliability by implementing automation, infrastructure as code, and monitoring to enhance service performance and availability.
Manage incident response by handling production issues, reducing downtime, and collaborating across engineering and DevOps teams.
Develop and troubleshoot large-scale distributed systems in both on-premises and cloud environments with a focus on scalability, fault tolerance, and cost efficiency.
Experience with Infrastructure as Code tools such as Terraform, Ansible, or Kubernetes.
Familiarity with monitoring and incident management tools like Datadog, Prometheus, PagerDuty, or Opsgenie.
Worked in production support and incident management for cloud-based distributed systems.
Work Experience Required: Not explicitly mentioned in the JD
Experienced in automating CI/CD pipelines and managing scalable cloud infrastructure with a focus on reliability and performance optimization.
Demonstrated ability to collaborate cross-functionally with engineering, security, and compliance teams in a fast-paced environment.
Skilled in using a range of SRE tools including monitoring, logging, tracing, and security tools indicative of hands-on operational expertise.