





Niche ML-evaluation skills reduce applicants despite metro location and mid-level experience amplifiers.
Specialized ML/LLM evaluation skills transfer across ML-heavy industries but less to non-tech sectors.
Explicit 5–8 years requirement plus specialized LLM evaluation, Python, and evaluation-platform experience.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and build Tekion’s AI evaluation platform and frameworks to enable measurement, trust, and improvement of AI outputs quality across multiple ML teams.
Design and implement evaluation datasets, automated scoring pipelines, and quality metrics for accuracy, consistency, safety among AI services covering pre-release benchmarking to continuous production monitoring.
Leverage AI and LLMs to automate evaluation scaling, including automated judges and synthetic dataset generation, and establish quality gates in CI/CD pipelines for model and data changes.
5–8 years experience in SDET, quality engineering, ML engineering, or data science with hands-on evaluation or measurement system building, or strong SDET background with LLM/ML fluency.
Strong programming skills in Python to build reusable evaluation tooling and pipelines.
Experience with ML/LLM evaluation methodologies, benchmark design, and addressing evaluation challenges of non-deterministic systems.
Work Experience Required: 5–8 years relevant experience. Notice Period: Not explicitly mentioned in the JD.
Experienced with LLM-as-judge, rubric-based scoring, or human-in-the-loop evaluation approaches within production AI/ML environments.
Demonstrates deep domain knowledge of LLM/agent concepts such as prompting, retrieval-augmented generation (RAG), embeddings, and generative failure modes (hallucination, bias, drift).
Skilled in designing and curating datasets, statistical interpretation of evaluation results, and collaborating cross-functionally to translate quality signals into actionable decisions.