





Metro senior niche AI-SRE role reduces general applicant density, moderate competition.
Core SRE skills are transferable, but AI/LLM operational experience increases domain specificity.
Explicit 8+ years, AI-production experience, and specific SRE/tooling requirements enforce strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end reliability and operational health of AI production systems including LLM services, RAG pipelines, and ML inference platforms.
Define and maintain SLOs/SLIs/SLAs and design high-availability architectures to achieve 99.9%+ uptime for AI products.
Lead incident management, build observability stacks, and automate detection/remediation of AI system issues.
8+ years professional experience in software engineering, DevOps, or site reliability engineering, including 1+ year operating AI/ML production systems.
Bachelor's degree in Computer Science, Software Engineering, or related field (or equivalent experience).
Proven experience maintaining high availability (99.9%+) and expertise with observability tools like Prometheus, Grafana, or Datadog.
Advanced experience with cloud platforms (GCP, AWS), Kubernetes, Infrastructure-as-Code (Terraform), and programming in Python plus one systems language.
Experienced in operational aspects of AI/ML systems, including managing LLM provider dependencies and AI-specific failure modes.
Strong leadership in incident response and operational risk management for complex AI workloads at scale.
Skilled in designing and improving CI/CD pipelines, infrastructure automation, and cloud-native deployment for AI platforms.