





Tier-1-backed Bengaluru role but niche ML/evals focus and seniority produce moderate competition.
Role demands specialized LLM/agent and evaluation expertise, limiting easy cross-industry transferability.
Requires concrete LLM/agent shipping experience, strong TypeScript/Python and ML evaluation skills, creating strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own AI quality for Cardboard’s agentic video editor by defining quality standards and building trusted evaluation datasets from real product usage.
Develop offline and online evaluation systems including automated checks, model graders, and human reviews, and track AI quality metrics like latency and cost.
Analyze agent failures, implement regression checks and release gates, and collaborate with product and engineering teams to deliver measurable AI quality improvements.
Experience shipping and operating an LLM or agent system with real customers.
Strong software engineering skills in TypeScript or Python with cross-language proficiency.
Experience in building evaluation datasets, experiments, or AI quality systems.
Work Experience Required: Not explicitly mentioned in the JD. Onsite Location: Bengaluru, India.
Senior individual contributor with strong ownership and product judgment able to translate vague AI quality issues into measurable problems.
Experience working across data, evaluation methods, model selection, and fine-tuning in applied ML contexts.
Bonus if experienced with multimodal AI, video/media applications, human labeling, model graders, or strong experiment design and statistical understanding.