Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Monitor and manage the health, latency, availability, and error rates of AI services and infrastructure, including incident triage and resolution with root-cause analysis and blameless post-incident reviews.
Build and maintain observability tools like dashboards, alerts, and logging for LLM workloads and improve deployment pipelines and automation for AI platform components.
Serve as escalation point for AI platform issues, maintain support documentation, and assist in onboarding new teams and clients onto the platform.
Minimum Requirements
3+ years experience in platform support, site reliability engineering, DevOps, or production support for cloud-hosted services.
Hands-on experience supporting applications built on LLM APIs such as OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, or Google Vertex AI.
Strong working knowledge of at least one major cloud platform (AWS, Azure, or GCP), container orchestration (Docker, Kubernetes), scripting in Python plus Bash or PowerShell.
Experience with observability tooling (Datadog, Grafana, Prometheus, CloudWatch, Azure Monitor), CI/CD and infrastructure-as-code tools, and incident management including post-incident report writing.
Ideal Candidate Profile
Experienced in managing production AI/ML platform reliability and observability with operational ownership of LLM-based services.
Proficient in modern cloud-native DevOps and SRE practices with a focus on automation, monitoring, and incident response for AI workloads.
Comfortable working in client-facing or consultancy environments, preferably with regulated sectors experience and capable of translating technical issues for non-technical stakeholders.
