






Mid-level metro role with common automation skills but niche AI-evaluation focus increases competition.
Requires AI/LLM evaluation experience plus QA automation, moderately limiting cross-industry transferability.
Explicit 3–6 years plus required SDET/ML evaluation and Playwright/Python skills enforces strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and develop evaluation harnesses and test infrastructure for AI-facing features to measure system quality across metrics like capture, retrieval, and guidance quality.
Manage end-to-end and API-level test frameworks supporting continuous daily releases, including Playwright-class tools.
Design and run host-behavior probes via scripted user sessions across various AI assistants to ensure consistent and expected product behavior, gating production releases with thorough UAT and regression analysis.
3–6 years of relevant experience in SDET/QA automation or ML evaluation with ownership of test or evaluation infrastructure.
Strong programming skills in Python or TypeScript, particularly in API-level testing with tools like Playwright or Cypress.
Experience or demonstrated interest in evaluating non-deterministic systems such as LLM applications.
Work Location: HSR, Bengaluru (On-site).
Proven operator in building and owning evaluation/test infrastructure for AI or ML systems rather than only manual test execution.
Experienced in working with diverse AI assistants and adapting evaluation methodologies to host-specific and evolving AI behaviors.
Autonomous worker capable of defining and managing their own evaluation and verification workflows to ensure product reliability at scale.