





Recognizable brand, metro location, and common 5+ experience level produce medium competition.
Specialized LLM evaluation and production ML experience is transferable but remains domain-sensitive.
Explicit 5+ years plus mandatory production ML, LLM, Python, and cloud experience enforces strict filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead AI evaluation and quality strategy for Generative AI systems including LLMs, RAG, and agentic workflows.
Design and implement scalable evaluation frameworks, automated validation pipelines, and observability standards integrated into CI/CD for production AI systems.
Contribute to AI-powered application development and ensure alignment with Responsible AI principles and enterprise governance.
Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Engineering, or equivalent practical experience.
5+ years of hands-on experience delivering AI/ML-powered systems into production.
Strong proficiency in Python and experience with AI evaluation, quality measurement, and observability in production environments.
Experience building cloud-native applications on AWS, Azure, or Google Cloud.
Experienced in leading AI evaluation and quality engineering for large-scale Generative AI including LLMs and RAG systems.
Skilled in designing automated testing and benchmarking frameworks with strong knowledge of AI observability and Responsible AI.
Proficient in collaborating cross-functionally with product and engineering teams in cloud-native AI product environments.