





Metro location, broad DevOps skillset, and a popular SRE role increase candidate competition.
Role requires specialized SRE/cloud skills that are somewhat transferable across industries.
Mandatory specific cloud, Kubernetes, IaC and monitoring tool experience increases screening stringency.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Manage and escalate system alerts to maintain high availability and participate in 24x7 on-call rotation for critical SaaS incidents.
Lead incident response, conduct root cause analysis, and enforce error budgets to balance development speed with system stability.
Automate operational tasks and design, deploy, and maintain cloud infrastructure using Terraform and Helm on EKS/Kubernetes clusters; maintain and evolve CI/CD pipelines and observability platforms like Datadog.
Hands-on AWS Cloud experience with strong knowledge of AWS IAM roles and policies.
Proficiency with EKS/Kubernetes and Infrastructure as Code tools: Terraform and Helm.
Experience with Docker, Docker Swarm, monitoring tools including Datadog, Grafana, Prometheus, and scripting in Bash/Python.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in managing high-availability SaaS products with a focus on reliability and automation in cloud environments.
Practical expertise in DevSecOps, incident management, and operational excellence within a 24x7 support and escalation context.
Skilled in designing observability strategies and proactive monitoring with customer impact awareness during deployments and updates.