





Mid-level, metro SRE lead with popular 5-8yr band but specialized AI/observability skills.
Core cloud and SRE skills are transferable, but AI/manufacturing reliability focus increases domain sensitivity.
Multiple mandatory technologies and explicit 5-8 years requirement increase shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead reliability engineering for key manufacturing and AI/Agentic-AI platforms including defining and enforcing SLOs, observability, CI/CD pipelines, infrastructure automation, and incident management.
Own AI/Agentic-AI reliability aspects including model and agent observability, drift monitoring, AIOps implementation, and GenAI-assisted incident response runbooks.
Drive production readiness through load/stress testing, on-call incident response, mentoring, and cross-functional collaboration with platform, manufacturing, data/AI and IT teams.
5-8+ years experience in SRE, DevOps, or platform engineering.
Strong expertise in SLI/SLO definition, error-budget management, incident management, and observability tooling (Splunk, Datadog, Prometheus, Grafana).
Proficient with GCP (networking, compute, managed services), Kubernetes, Docker, Terraform, Jenkins CI/CD; GitOps/Argo CD exposure desirable.
Experience in AI/ML reliability, MLOps/AIOps, including model monitoring, drift detection, and GenAI-assisted incident response.
Experienced in leading technical reliability and architectural decisions across distributed and AI systems with measurable impact on MTTA/MTTR and operational toil reduction.
Skillful in designing and operating integrated observability and AIOps frameworks supporting both manufacturing and AI-driven platforms.
Comfortable with flexible onsite collaboration, incident ownership, and mentoring within complex cross-functional teams.