





Mid-level metro role with moderate specialization and no strong brand, so medium competition.
LLM evaluation focus and non-deterministic testing create high domain-specificity.
Explicit 3–6 years plus mandatory SDET automation and Playwright/Python skills, so high strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the evaluation scaffolding and quality layer for AI-facing features, including test infrastructure and release gates.
Build and maintain evaluation harnesses and execute host-behavior probes across diverse AI assistants like Claude and ChatGPT.
Gate production releases through user acceptance testing, regression analysis, and quality reporting to ensure system reliability.
3-6 years of experience in SDET/QA automation or ML evaluation with ownership of test or evaluation infrastructure.
Strong programming skills in Python or TypeScript.
Experience with API-level testing and tools like Playwright or Cypress.
Location: Bengaluru (Hybrid); Start: In 2–3 months.
Experienced in building and owning test/evaluation infrastructure, not just executing tests.
Able to work autonomously to define workflows for evaluation and verification.
Familiarity or strong interest in LLM applications and evaluating non-deterministic AI systems.