






Niche LLM-eval skillset lowers competition despite mid-level experience requirement.
Specialized LLM evaluation skills transferable across industries but require ML/LLM expertise.
Explicit 4+ years plus mandatory LLM-eval tooling, Python, and CI make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and build the evaluation infrastructure for AI outputs, including datasets, scorers, harnesses, and CI gates to enforce quality before shipping.
Define and enforce 'good enough to ship' criteria by measuring grounding, faithfulness, hallucination, and correctness of multi-step agentic AI flows.
Create reproducible evaluation systems that block regressions and provide trusted metrics to ensure AI release quality and continuous monitoring.
4+ years in software or ML engineering with hands-on experience building LLM evaluation or quality tooling.
Strong Python programming skills and experience with CI/CD processes (e.g., GitHub Actions).
Proven knowledge of grounding, faithfulness, hallucination concepts and how to measure them rigorously in AI outputs.
Familiarity with LLM evaluation frameworks and LLM-as-judge methods for scoring model outputs.
Experienced in designing robust, reproducible evaluation metrics and CI gates for non-deterministic LLM/agentic systems to prevent quality regressions.
Skilled at building and curating adversarial and edge-case labeled datasets targeting correctness-critical outputs.
Preferably has domain exposure to high-stakes environments (e.g., FinTech) or multi-step agentic AI evaluation workflows.