





Mid-level, metro location, and a strong employer increase competition, despite niche ML research skills.
Specialized LLM research and evaluation expertise is tightly domain-specific and less transferable.
Explicit 5+ years requirement plus mandatory LLM training, PyTorch, and agentic research experience.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and lead research on large language model (LLM) mid-training and post-training processes, including data mixture decisions and impact on agent behavior.
Research and prototype advanced agentic AI architectures covering planning, reasoning, memory, skills, tool use, retrieval, and multi-agent collaboration, pushing beyond current methods.
Design and develop systematic experimentation frameworks, novel evaluation methodologies, and benchmarks to rigorously measure agent abilities such as reasoning, planning, reliability, and safety.
5+ years of professional experience in machine learning, deep learning, AI research, or related applied research roles.
Hands-on experience with LLM training and post-training techniques including continued pretraining, supervised fine-tuning (SFT), preference optimization, or reinforcement learning (RL).
Expertise in Python and advanced PyTorch, with experience modifying models and training infrastructure for experimental research.
Experience in designing evaluation methodologies and benchmarks, including LLM-as-a-Judge and human evaluation approaches.
Strong background in agentic AI with practical experience in planning, reasoning, memory, long-context processing, tool use, and knowledge grounding.
Proven research ownership demonstrating ability to independently identify research problems, formulate hypotheses, and drive projects to measurable impact and production transition.
Skilled at defining rigorous evaluation strategies and translating research insights into improved training data, architectures, and evaluation metrics.