





Metro Bengaluru location and broad AI+QA skillset raise competition, niche specialization reduces applicant volume.
Role requires specialized LLM, RAG, and AI-observability skills, limiting cross-industry transferability.
Mandatory 7+ years and specific Python, Pytest, Playwright, LLM evaluation and observability skills enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end quality assurance for enterprise-grade AI applications including LLMs, RAG systems, and AI agents.
Design and implement automated testing pipelines, benchmark datasets, and quality gates focusing on accuracy, reliability, performance, and scalability.
Collaborate across AI Engineering, Product, and Platform teams to validate AI workflows, monitor production readiness, and conduct root cause analysis using observability tools.
7+ years of experience in Software QA, Test Automation, or AI Quality Engineering.
Strong Python programming skills with hands-on experience using Pytest for unit, integration, API, regression, and end-to-end testing.
Experience evaluating LLM applications and Retrieval-Augmented Generation (RAG) systems with quality metrics like hallucination rate, tool selection accuracy, precision/recall, and latency.
Experience with frontend automation using Playwright and observability tools like OpenTelemetry integrated into CI/CD pipelines.
Experienced in AI quality engineering specifically with large language models, AI agents, and RAG workflows showing deep technical expertise.
Skilled at designing measurable quality standards and automated evaluation benchmarks operationalized in CI/CD pipelines.
Proficient in debugging, root cause analysis, and performance validation within complex AI production environments.