Senior Site Reliability Engineer
METRO Global Solution Center IndiaMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Ensure stability and reliability of GCP-based cloud-native applications using Kubernetes and Docker.
Define and monitor SLOs/SLAs/SLIs; develop and maintain monitoring and alerting systems with Datadog and GCP tools.
Automate infrastructure provisioning and manage Kubernetes configurations using Terraform, Helm, and Kustomize; conduct incident analysis and collaborate on CI/CD integration (GitHub Actions).
Minimum Requirements
5+ years of Site Reliability Engineering experience.
Strong expertise in Google Cloud Platform, Kubernetes, Docker, Terraform, Helm, and Kustomize.
Experience with monitoring tools like Datadog and GCP Monitoring; scripting skills in bash and automation frameworks.
Bachelor’s or Master’s degree in Computer Science, Software Engineering, or equivalent practical experience.
Ideal Candidate Profile
Experienced in designing and operating scalable and resilient systems in cloud environments, especially GCP.
Proficient in end-to-end reliability practices including infrastructure automation, observability, and incident response.
Operates effectively both independently and collaboratively in fast-paced, Agile/Scrum environments.
