Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Lead testing and quality assurance for AI-powered applications including Generative AI, Retrieval Augmented Generation (RAG) systems, AI agents, and Large Language Model (LLM) based solutions.
Design and execute evaluation frameworks using tools like DeepEval, RAGAS, LangSmith, and develop automated AI evaluation pipelines integrated with CI/CD.
Collaborate with cross-functional teams to define test strategies and validate AI-enabled features, ensuring reliability, safety, and robustness across multiple business scenarios.
Minimum Requirements
Hands-on experience testing AI/GenAI applications in production or pre-production environments.
Strong understanding of Large Language Models, Generative AI concepts, prompt engineering, embeddings, vector databases, and RAG architectures.
Experience with AI evaluation frameworks such as DeepEval, RAGAS, LangSmith, or similar tools.
Work Experience Required: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced in testing AI agents and agentic frameworks like LangGraph, CrewAI, AutoGen, or Semantic Kernel.
Familiarity with AI security risks such as prompt injection, jailbreak attempts, data leakage, and model misuse.
Able to build and maintain automated AI evaluation frameworks and understand AI observability, monitoring, and quality metrics including hallucination detection and factual correctness.
