





Hybrid metro role with popular DevOps title but niche LLM requirements creates moderate applicant competition.
Core SRE skills transfer well, but hands-on LLM and eval experience increases domain specificity.
Requires mandatory production SRE experience, Python, Kubernetes, and hands-on LLM work.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and maintain the LLM gateway that routes and controls all AI model provider access, ensuring high availability since all AI features rely on it.
Manage runtime reliability with full SRE responsibilities including SLOs, on-call, incident response, and postmortems for production systems.
Control LLM-related costs with per-team/model budgets and quotas, while building observability, caching, guardrails, and evaluation infrastructure for AI runtime services.
Experience operating production infrastructure with Kubernetes, AWS, Terraform, and GitOps, including on-call and incident management experience.
Proficiency in Python for owning services and tooling beyond scripting.
Hands-on experience building and using LLM-based tools or pipelines (agents, RAG pipelines, internal tools) with knowledge of tokens, context windows, prompt/response tracing, and evaluation frameworks.
Fluency in observability including instrumentation, metrics/log analysis, alerting beyond just dashboard consumption.
Background in production-grade infrastructure and SRE practices supporting AI runtime systems with responsiveness to incidents and outages.
Operational familiarity with modern LLM gateways, observability tooling, cost attribution (FinOps), and guardrail enforcement in AI applications.
A pragmatic practitioner who has built and improved AI feature infrastructure spanning model routing, caching, evaluation, and cost control in a dynamic environment.