





Niche senior AI-safety and Google ADK requirements greatly reduce candidate competition.
Highly specialized ML/AI safety and LLM evaluation skills limit cross-industry transferability.
Mandatory 8+ years, safety Evals, Google ADK familiarity, and Python make screening strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Define, develop, and execute evaluation frameworks and test strategies for AI agents built with Google Agent Development Kit (ADK).
Ensure AI agents meet standards for safety, ethics, and high-quality performance with a focus on Responsible AI and Safety Evaluations.
Perform approximately 70% automation and 30% manual testing in evaluation processes.
8+ years of Software QA experience, including 2-3 years testing or evaluating AI/ML systems, conversational agents, or Large Language Models.
Mandatory experience in safety evaluations including red teaming, adversarial testing, bias detection, and toxicity measurement in generative AI.
Strong proficiency in Python programming for scripting, data processing, and automation; familiarity with PyTest is mandatory.
Direct or strong conceptual knowledge of Google Agent Development Kit (ADK); familiarity with Google Cloud Platform (e.g., Vertex AI) and MLOps integration.
Experienced in designing and executing safety evaluation workflows for AI agents ensuring ethical and safe deployment.
Skilled in bridging AI development and deployment by integrating automation and manual testing practices.
Practically familiar with AI evaluation tools and frameworks such as Langsmith, DeepEval, Ragas, Giskard, and Hugging Face.