





Metro location, popular SRE role, and broad skill requirements increase applicant competition.
SRE and observability skills are transferable across industries but require significant domain-specific experience.
Multiple mandatory SRE, observability, IaC and cloud requirements create high technical filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead integration of AI/ML technologies like LLMs into the observability platform for automated root-cause analysis, anomaly detection, and self-healing.
Own and ensure performance, scalability, availability, and cost-efficiency of the global observability platform and related automation using IaC tools such as Terraform.
Advise engineering teams on best practices for product/platform reliability, manage data quality standards, and support problem management impacting customer experience.
Experience with AI/ML integration in operational workflows (AIOps) and proficiency in monitoring/logging/tracing tools like NewRelic, Datadog, or Splunk.
Hands-on expertise with OpenTelemetry framework and strong coding skills in Python, Go, or Java; familiarity with shell scripting.
Proficiency with Infrastructure as Code tools such as Terraform, Ansible, or Puppet and experience with major cloud providers like AWS, GCP, or Azure.
Work Experience Required: Not explicitly mentioned in the JD
Experienced in embedding AI-driven automation in large-scale observability platforms within global organizations.
Strong operator with a deep understanding of SRE methodologies (Google’s Golden Signals, RED, USE) and platform scalability.
Proven ability to lead cross-functional engineering and site reliability teams to operational excellence and customer-focused outcomes.