






Tier‑1/YC backing, generalist applied-ML role, and Bangalore metro drive high competition.
Requires LLM/agent and AI-product experience, moderately specialized but transferable across ML roles.
Requires demonstrable LLM/agent shipping, evaluation systems, and engineering skills, so moderately strict screening.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the measurement and improvement of Cardboard’s AI agent quality through building evaluation datasets and tracking metrics.
Analyze model and agent failures to enhance quality via data, evaluation methods, model selection, and fine-tuning where applicable.
Implement regression checks and release gates for quality, latency, and cost while collaborating with product and engineering teams to ship improvements.
Experience shipping and operating large language model (LLM) or agent systems with real customers.
Strong software engineering skills in TypeScript or Python, able to work across both.
Experience building evaluations, datasets, experiments, or AI quality systems.
Work Experience Required: Not explicitly mentioned in the JD.
Has demonstrated ability to convert ambiguous AI quality problems into measurable improvements with strong product judgment.
Experienced working end-to-end on reliable AI products rather than only research or academic work.
Comfortable working in a fast-moving small team environment alongside technical founders and engineers.