Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Own and build the evaluation infrastructure for AI outputs, including datasets, scorers, harnesses, and CI gates to enforce quality before shipping.
Define and enforce 'good enough to ship' criteria by measuring grounding, faithfulness, hallucination, and correctness of multi-step agentic AI flows.
Create reproducible evaluation systems that block regressions and provide trusted metrics to ensure AI release quality and continuous monitoring.
Minimum Requirements
4+ years in software or ML engineering with hands-on experience building LLM evaluation or quality tooling.
Strong Python programming skills and experience with CI/CD processes (e.g., GitHub Actions).
Proven knowledge of grounding, faithfulness, hallucination concepts and how to measure them rigorously in AI outputs.
Familiarity with LLM evaluation frameworks and LLM-as-judge methods for scoring model outputs.
Ideal Candidate Profile
Experienced in designing robust, reproducible evaluation metrics and CI gates for non-deterministic LLM/agentic systems to prevent quality regressions.
Skilled at building and curating adversarial and edge-case labeled datasets targeting correctness-critical outputs.
Preferably has domain exposure to high-stakes environments (e.g., FinTech) or multi-step agentic AI evaluation workflows.

