





Reputable startup brand, mid-level 3–5 years, and Bengaluru metro location increase competition.
Requires AI/NLP evaluation and labeling expertise, so moderately sensitive to industry background.
Explicit 3–5 year requirement plus mandatory AI-eval and SQL skills make filters moderately strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the human evaluation system to assess and ensure AI product quality through labeling, categorizing failure modes, and maintaining benchmark data.
Analyze AI output quality issues (hallucinations, retrieval failures, tool-use problems) and translate ambiguous quality issues into structured, actionable feedback.
Collaborate cross-functionally with QA, Engineering, Product, and eval teams on quality reporting, benchmark updates, and release readiness decisions.
3–5 years of experience in data labeling, data analysis, QA, or related fields.
Experience evaluating AI-generated outputs or working with NLP, search, recommendation, or other ML systems.
Familiarity with structured qualitative analysis, basic SQL/data retrieval, and spreadsheet workflows.
Hybrid work location requirement (4 days/week in office).
Detail-oriented analyst skilled at applying nuanced rubrics consistently for human AI output evaluation.
Experienced in handling multi-step AI workflow evaluations and interpreting qualitative failure patterns for product improvement.
Comfortable working independently in fast-moving, ambiguous environments and partnering cross-functionally to drive measurable quality improvements.