





High due to Tier-1 brand, metro location, and mid-level generalist ML/AI role with broad requirements.
Medium — ML engineering skills are transferable, but forensic/LLM benchmarking context adds domain specificity.
High because the JD mandates explicit years plus specific ML, cloud, CI/CD, and LLM evaluation skills.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain AI benchmarking and experimentation platforms to support model evaluation and client use cases.
Design end-to-end benchmarking workflows and build scalable evaluation frameworks integrating engineering and academic research.
Produce client-ready insights, support technical demos, and mentor junior team members while deepening technical skills and client understanding.
Bachelor's degree required.
3-7 years of relevant work experience.
Proficiency in English (oral and written) is mandatory.
Experience with Python programming, cloud platforms (Azure, AWS, GCP), CI/CD processes, and containerization tools like Docker or Podman.
Experience applying statistics, experimental design, and LLM evaluation techniques including LLM-as-a-Judge.
Skilled in building scalable AI research infrastructure and benchmarking frameworks within forensic technology contexts.
Able to navigate complex client scenarios, contribute to technical strategy, and thrive in collaborative, fast-paced environments.