





Strong employer brand, remote option, metro location, and broad technical requirements increase candidate competition.
Role requires deep SRE-specific skills and tooling, limiting cross-domain transferability.
Extensive mandatory SRE tooling, IaC, Kubernetes and observability requirements increase screening rigidity.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead reliability engineering efforts to improve system fault tolerance, monitoring, and incident response to meet business SLAs and SLOs.
Own incident management lifecycle including prioritization, triage, communication, mitigation, post-mortem analysis, and corrective actions.
Manage stakeholder expectations during incidents, act as technical liaison with client engineering and executives, and mentor other SREs.
Proficient programming skills in one or more high-level languages (e.g., Python, Golang, Shell, Ruby, Java).
Expertise with DevOps/GitOps tools and CI/CD pipeline integration (e.g., GitLab, Jenkins, CircleCI).
Experience with IaC tools like Terraform, Ansible, ARM, CloudFormation.
Work Experience Required: Not explicitly mentioned in the JD.
Deep knowledge of observability tooling (Grafana, Prometheus, ELK, Jaeger, etc.) and container orchestration (Kubernetes, AWS EKS, Docker Swarm).
Experience driving system performance tuning, incident management, and collaboration with cross-functional teams under pressure.
Strong communicator comfortable interfacing with both technical teams and C-level executives on reliability and incident resolution.