Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and develop the evaluation platform pipelines that run AI agent traces through model-based judges at scale.
Design and implement evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety.
Take research problems from concept to prototype to shipped feature, ensuring evaluation results are actionable for teams.
Minimum Requirements
5+ years experience in machine learning, applied AI, prompt engineering, or agentic AI with real user-facing product delivery.
Strong proficiency in Python programming.
Practical experience with agentic AI concepts: planning, reasoning, memory, tool use, retrieval, and long-context handling.
Experience designing evaluation methodologies, not just conducting evaluations.
Ideal Candidate Profile
Demonstrated end-to-end ownership of AI evaluation systems and pipelines, including scoring logic and data curation.
Hands-on expertise with LLM APIs and prompt engineering for production environments, addressing cost and latency tradeoffs.
Experience communicating technical concepts effectively to both technical and non-technical stakeholders.
