





Strong Tier-1 brand, remote role, mid-level SRE, and generalist toolset increase applicant competition significantly.
Strong SRE and observability focus is transferable across industries but requires security-first mindset.
Explicit 5+ years requirement plus extensive mandatory SRE, observability, and CI/CD tool expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain a scalable, highly reliable internal developer platform to accelerate software delivery.
Build automation solutions and monitoring tools to improve system reliability, performance, and serviceability.
Lead incident response, production readiness reviews, and integrate AI-enabled observability and self-healing capabilities into platform operations.
5+ years experience in large-scale production environments with CI/CD, IaC, monitoring, and Kubernetes technologies.
Strong hands-on skills with CI/CD tools (Bazel, GitHub Actions, Jenkins), IaC (Ansible, Terraform, Puppet), and observability tools (Prometheus, Datadog, Grafana, etc).
Proficiency in programming languages such as Python, Go, TypeScript, Java, or C#.
Work Experience Required: 5+ years in relevant SRE or site reliability domains.
Experienced in operating at enterprise scale with both on-premise and cloud infrastructure for developer platforms.
Skilled in integrating AI and automation into observability and incident management workflows to achieve proactive reliability and zero-toil operations.
Comfortable mentoring peers and collaborating with distributed engineering teams while managing trade-offs between short-term fixes and long-term reliability goals.