





Mid-level, niche LLM evaluation role at a known startup leads to moderate applicant competition.
Specialized LLM evaluation skills limit transferability across industries.
Explicit 5–8 year requirement plus mandatory ML/LLM and Python skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the design and development of Tekion’s AI evaluation platform and infrastructure to measure quality of AI outputs across various business domains.
Build and maintain evaluation datasets, automated scoring pipelines, and quality metrics for accuracy, consistency, safety, and task success of AI models and agents.
Implement offline and online evaluation systems including benchmarks, continuous quality monitoring, evaluation gates in CI/CD, and dashboards for actionable AI quality insights.
5–8 years experience in SDET, quality engineering, ML engineering, or data science with hands-on experience building evaluation or measurement systems, or strong SDET background with LLM/ML fluency.
Proficient in Python programming for building evaluation pipelines and tooling.
Experience in ML/LLM evaluation methods including LLM-as-judge, rubric-based scoring, and human-in-the-loop evaluation.
Deep understanding of LLM/agent concepts, generative failure modes (hallucination, drift, bias), and designing/curating datasets with labeling and quality control.
Experienced in building shared AI evaluation platforms or capabilities used across multiple ML teams.
Strong statistical intuition and skill in translating evaluation results into quality decisions for ML and product teams.
Comfortable working closely with ML engineers, data scientists, and product management to align evaluation strategies with business domain needs.