





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
High due to strong employer brand, metro location, and a visible mid-level generalist support title.
Medium because cloud, observability, and SQL skills transfer broadly, while LLM-specific troubleshooting remains specialized.
High because multiple explicit mandatory years and technical stack/LLM/cloud requirements narrow candidate pool.
Drive incident resolution, root-cause analysis, and performance optimization for Workday’s AI/ML enterprise platform and autonomous AI agent workflows.
Analyze system metrics, debug cloud-based ML service pipelines and LLM orchestration layers, and manage critical customer escalations within SLAs.
Partner with engineering and data science teams to improve AI prompt, data, and workflow based on evaluation of LLM outputs and system traces.
Minimum 3 years experience in technical support engineering, platform operations, or escalation management for enterprise SaaS platforms.
At least 2 years hands-on experience with Large Language Model (LLM) pipeline analysis or troubleshooting and cloud diagnostics on AWS, GCP, or Azure.
Minimum 2 years experience using monitoring tools (Grafana, Kibana, Datadog, Prometheus) and programming or SQL for data analysis.
Willingness to participate in weekend on-call rotations. Notice period: Not explicitly mentioned in the JD.
Experienced in high-severity incident management and customer escalation in enterprise SaaS and AI platforms.
Strong technical troubleshooting skills across cloud infrastructure, AI/ML workflows, and backend data validation.
Ability to effectively communicate complex technical issues to cross-functional teams and non-technical stakeholders.