





Niche AI-assurance skills and specific eval tools limit applicant pool despite metro location.
Highly specialized AI-assurance and regulatory knowledge limits industry transferability.
Multiple mandatory ML-assurance skills, explicit 5–7 years, and platform/regulatory requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and maintain evaluation datasets and scenarios for 19 Trustworthiness checks across Gemini-based agents.
Run scheduled and on-commit evaluation suites (bias, toxicity, red-team, explainability), triage failures to appropriate owners.
Maintain CI/CD threshold gates blocking release on failing checks and package audit evidence for compliance reviews.
5-7 years of QA/test engineering experience for ML or GenAI systems with dedicated evaluation frameworks (e.g., DeepEval, Ragas, Promptfoo).
Hands-on experience with AssureAI or comparable AI-assurance/evaluation platforms.
Proficient in Python and experience integrating test suites into CI/CD pipelines (e.g., Cloud Build, GitHub Actions).
Familiarity with AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) and translating them into test coverage.
Experienced in bias/fairness testing, explainability techniques (SHAP/LIME), and adversarial prompt/red-teaming methodologies.
Detail-oriented operationally handling multiple failing checks across agents and scaling evaluation throughput.
Capacity to produce audit-evidence packages regularly and proactively track and close dataset coverage gaps as new agents are onboarded.