





Tier-1 brand, metro location, and broad technical skill requirements increase candidate density.
Role requires specialized LLM evaluation and AI governance expertise, limiting cross-industry transferability.
Multiple mandatory niche skills (LLM evaluation, cloud, CI/CD, governance) imply strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design and execution of comprehensive test strategies for AI/ML solutions focusing on accuracy, bias, robustness, and regression.
Develop and implement rubric-based evaluation frameworks and automated validation suites for GenAI systems, integrating with CI/CD pipelines.
Partner cross-functionally to ensure AI data quality, security, compliance, and continuous iteration of evaluation methodologies.
Proficiency in programming languages such as Python or TypeScript used in AI/test automation.
Experience with rubric-based evaluation and LLM evaluation frameworks or benchmarking tools (e.g., LangSmith, Confident AI).
Knowledge of GenAI solutions including Retrieval Augmented Generation and cloud AI platforms (Azure OpenAI, Vertex AI, ChatGPT Enterprise).
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in designing and operationalizing AI/ML testing strategies with a focus on GenAI and large language models.
Comfortable collaborating across IT, engineering, compliance, and risk teams to enforce AI governance and ethical standards.
Detail-oriented analytical thinker capable of detecting subtle AI model issues like hallucinations and documenting findings precisely.