





Niche AI QA skillset reduces applicant pool despite Bengaluru location.
Strong AI-specific evaluation and tooling requirements limit transferability across industries.
Mandatory 7+ years and specific AI QA tooling, Python, Pytest, and benchmarks make filtering strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end quality assurance and test automation for AI-powered applications including LLM-powered and conversational AI systems.
Design and implement benchmarking methodologies and create golden datasets for AI regression testing.
Evaluate AI system performance using industry-standard metrics and AI observability platforms to ensure reliability and accuracy.
7+ years of experience in Software Quality Assurance, Test Automation, or AI Quality Engineering.
Proficient in Python with experience using Pytest for automated testing (unit, integration, API testing).
Hands-on experience evaluating AI systems including LLMs, conversational AI, NL2SQL, and understanding AI benchmarking.
Strong knowledge of AI evaluation metrics (Precision, Recall, F1, EM, MRR, latency), SQL skills for query validation, and experience with AI evaluation platforms like Lang Smith, MLflow, or Arize Phoenix.
Experience working closely with AI teams to build automated, scalable QA frameworks tailored to AI workflows.
Ability to conduct detailed root cause analysis and debugging across distributed AI systems and backend services.
Deep understanding of test automation best practices combined with AI domain expertise to drive quality improvements in generative AI or conversational AI products.