





Mid-level metro SRE role with common cloud and Kubernetes skills results in medium competition.
Requires SRE, cloud and AI-model reliability expertise, making background fit highly domain-specific.
Explicit years plus mandatory GCP, Kubernetes, Terraform, observability and leadership requirements create high strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead reliability engineering efforts for critical manufacturing and AI/Agentic-AI platforms, owning SLOs, observability, CI/CD pipelines, infrastructure automation, incident management, and production readiness.
Define and enforce SLI/SLO/SLA targets including error budgets, developing dashboards and alerts to optimize operational performance and reduce alert fatigue.
Implement and oversee AI-specific reliability practices such as model/agent monitoring, AIOps for anomaly detection and GenAI-assisted incident management.
5-8+ years of experience in Site Reliability Engineering, DevOps, or platform engineering.
Strong expertise in SLI/SLO/error budget management, observability tools (Splunk, Datadog, Prometheus, Grafana), incident management, and CI/CD (Jenkins mandatory, GitOps/Argo CD desirable).
Hands-on production experience with GCP including networking and managed services, Kubernetes, Docker, Terraform, and paging/on-call tools (PagerDuty or equivalent).
Proficiency in Python and/or Java scripting; experience with load/stress testing, capacity planning, and production readiness.
Experienced technical leader with capability in architectural decision-making and cross-functional collaboration across platform, manufacturing, data/AI, and IT teams.
Prior exposure to AI/ML reliability, MLOps, AIOps, or Agentic-AI including Vertex AI, LangChain or comparable frameworks, and GenAI-assisted incident response workflows.
Track record of owning end-to-end incident management processes with metrics-driven improvement (e.g., MTTA/MTTR) and reducing operational toil through automation.