





Tier-1 employer, mid-level AI title, and 5+ years requirement increase applicant density.
Requires production ML/LLM evaluation expertise so skills are moderately transferable across industries.
Explicit 5+ years, production ML experience, LLM evaluation, Python, and cloud deployment make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead AI evaluation and quality strategy for Generative AI and ML systems, including defining and tracking key AI quality metrics like accuracy, bias, safety, and hallucination rates.
Design and implement scalable automated evaluation and validation frameworks for Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) systems, and agentic workflows integrated into CI/CD pipelines.
Collaborate across teams to build and deploy cloud-native AI-powered applications and ensure observability, reliability, and Responsible AI governance in production environments.
Bachelor's or Master's degree in Computer Science, Machine Learning, Data Science, Engineering, or equivalent practical experience.
Minimum 5 years of hands-on experience delivering AI and ML-powered systems in production environments.
Proficient in Python with experience building automation, evaluation frameworks, and testing infrastructure for AI models.
Experience deploying and operating AI applications on cloud platforms such as AWS, Azure, or GCP, including monitoring and continuous improvement.
Deep expertise in AI/ML evaluation methodologies and tooling, with hands-on experience in LLMs, RAG systems, and AI observability concepts.
Ability to lead technical projects that connect AI model evaluation to broader product and business goals, driving improvements in AI quality and engineering standards.
Experience working in cross-functional teams with senior technical leaders and a focus on operational excellence and Responsible AI principles.