





Strong employer brand, mid-level generalist SRE role, metro location, and broad cloud/Kubernetes requirements.
Core SRE skills (cloud, Kubernetes, IaC) are highly transferable across industries.
Explicit 3–7 years requirement plus mandatory cloud, Kubernetes, IaC, and on-call expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Ensure availability, reliability, scalability, and performance of SaaS production environments through monitoring, incident response, and on-call support.
Own cloud infrastructure management on AWS and/or GCP, Kubernetes cluster management (EKS/GKE), network, security, and cost optimization.
Automate operational tasks including CI/CD deployments, infrastructure as code (Terraform, Ansible), scripting, and documentation for continuous platform improvement.
3 to 7 years of relevant work experience in Site Reliability Engineering or production support roles.
Hands-on experience with cloud platforms (AWS/GCP), Kubernetes (EKS/GKE), infrastructure as code (Terraform, Ansible), and automation scripting (Python, Bash).
Experience with monitoring and incident management tools such as Datadog, PagerDuty, Prometheus, and Grafana.
Ability to respond to Priority-1 alerts promptly (acknowledge within 5 minutes, response within 15 minutes), including participation in on-call rotations.
Experienced in production incident response, root cause analysis, and continuous operational improvements in SaaS environments.
Comfortable operating in Agile, DevOps, or SRE teams with strong knowledge of SLI, SLO, and SLA metrics.
Skilled in managing Kubernetes and cloud-native services, with an emphasis on automation, cost optimization, and reliability engineering.