





Specialized senior SRE at a known employer in a metro increases competition moderately.
SRE and cloud skills are transferable, though AI/AIOps and OpenTelemetry specialization raises domain bias.
Extensive mandatory SRE, observability, IaC, cloud and AI/AIOps requirements make filters highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead integration of AI/ML capabilities into the observability platform to enable automated root-cause analysis, anomaly detection, and self-healing.
Maintain and optimize the global observability platform ensuring high availability, performance, and cost efficiency, including managing telemetry data quality and governance.
Develop automation and Infrastructure as Code tooling (e.g., Terraform) to streamline SRE tasks and support engineering teams in feature reliability and scalability.
Experience with AI/ML integration in operational workflows and familiarity with Large Language Models and prompt engineering.
Proficiency in monitoring and observability tools such as NewRelic, Datadog, Splunk, and hands-on experience with OpenTelemetry framework.
Strong coding skills in at least one language among Python, Go, or Java, plus experience with Infrastructure as Code tools like Terraform or Ansible.
Work Experience Required: Not explicitly mentioned in the JD. Location: Must be based in Chennai, India, with 4 days per week onsite availability. No travel required.
Experienced in designing and maintaining large-scale observability platforms with focus on reliability, scalability, and cost optimization.
Skilled in applying SRE frameworks and methodologies such as Google’s Golden Signals, RED, and USE methods to monitoring and alerting.
Strong operational mindset with expertise automating infrastructure and workflows for a global technology organization, particularly with cloud platforms (AWS/GCP/Azure), containerization (Docker), and orchestration (Kubernetes).