





Company brand, metro location, and broad tool requirements drive high competition among applicants.
SRE skills transfer across industries but require specific cloud, observability, and automation tool experience.
Explicit 10+ years requirement and mandated observability, cloud, and automation tooling indicate high shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain observability tools and infrastructure, including dashboards and alerts for system health.
Collaborate with engineering and product teams to enhance system reliability, improve incident response, and automate operational tasks.
Participate in system design consulting, capacity planning, incident response, and guide junior team members in best Site Reliability Engineering practices.
Engineering degree or related technical discipline with 10+ years of SRE experience.
Proficient in scripting languages such as Python, Bash, or Powershell and coding in higher-level languages like Python, JavaScript, C++, or Java.
Hands-on experience in Engineering or Cloud with knowledge of public cloud platforms (GCP, AWS, Azure), containerization, Terraform, Ansible, and CI/CD pipeline management.
Experience with observability tools including Splunk/ELK, Datadog, Prometheus/InfluxDB, Grafana, and distributed system design and architecture.
Experienced senior SRE with strong coding and infrastructure automation skills familiar with cloud-native environments and container technologies.
Proactive in driving technical and operational excellence by analyzing current technologies and implementing best practices in metrics, logging, and tracing.
Able to lead and mentor teams, collaborate cross-functionally, and influence stakeholders for optimal technical and business outcomes.