Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Drive reliability and scalability of production systems by maintaining health and service standards through monitoring, incident management, and automation.
Own and improve observability systems to provide actionable alerting and visibility for engineering teams.
Develop and manage Infrastructure as Code automation with tools like Terraform, Ansible, Kubernetes, and CI/CD pipelines to reduce operational toil and embed SRE practices into SDLC.
Minimum Requirements
5+ years of experience in Site Reliability Engineering, Software Engineering, or similar role.
Proficiency in programming/scripting languages such as Typescript, Java, Python, Go, PHP, or Ruby.
Experience with Unix/Linux administration, containerization (Docker, Kubernetes), cloud platforms (AWS, Azure, or GCP), and CI/CD tools (e.g., GitHub Actions, Argo CD).
Experience with relational databases like PostgreSQL, MySQL, or SQL Server.
Ideal Candidate Profile
Demonstrated ability to integrate AI-driven tooling and automation into SRE workflows and codebases for maintainability and resilience.
Experienced in incident command, root cause analysis, and embedding reliability practices such as SLIs/SLOs within product teams.
Comfortable collaborating cross-functionally with security and engineering to meet compliance standards and operational excellence in an AI-native SDLC environment.
