Senior Research Scientist, Agent Evaluation
ServiceNowMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and develop core evaluation platform components: pipelines that process agent traces and scoring logic converting outputs into actionable results.
Design and implement evaluation methodologies and benchmarks focusing on agent reasoning, planning, tool use, reliability, and safety across multiple evaluation types (LLM-as-Judge, trajectory-based, human).
Manage end-to-end product lifecycle from research question through prototype to shipped feature including dataset curation and evaluator performance measurement.
Minimum Requirements
5+ years experience in machine learning, applied AI, prompt engineering, or agentic AI with proven delivery of user-facing products.
Strong Python programming skills.
Practical experience in agentic AI domains such as planning, reasoning, memory, tool use, retrieval, and long-context handling.
Experience in designing evaluation methodologies beyond just running evaluations.
Ideal Candidate Profile
Expertise in building scalable ML pipelines and scoring systems for AI agent evaluation in production contexts.
Experience working hands-on with LLM APIs including prompt engineering, structured output formatting, and optimizing cost and latency.
Ability to translate complex technical results into clear insights for both technical and non-technical stakeholders.
