





Specialized LLM memory skills narrow applicants but mid-level ML demand keeps competition medium.
Highly specialized LLM memory, evaluation, and Bedrock experience limits transferability across generalist roles.
Explicit 4+ years, 2+ years LLM production, and AWS/metrics requirements create strict hiring filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain LLM-powered pipelines that convert raw agent interactions into structured memory, managing what is remembered, consolidated, retrieved, or forgotten.
Own and build evaluation systems to measure memory quality and its impact on agent performance, including benchmarks and A/B testing.
Collaborate to design memory extraction, consolidation, context assembly, and embedding strategies using LLMs and small tuned models within AWS ML infrastructure.
4+ years software engineering experience with strong Python skills, including 2+ years building production LLM or NLP applications.
Experience with prompt engineering, structured output extraction, RAG pipelines, embedding models, and building evaluation frameworks for LLM systems.
Solid ML foundations in retrieval metrics, ranking, classification, and statistics for A/B testing.
Working knowledge of AWS ML stack (Bedrock, SageMaker) or equivalent cloud ML platforms.
Experienced in designing and productionizing intelligent memory systems for LLM-powered agents, with emphasis on retrieval quality and memory consolidation.
Skilled in evaluation framework development for LLM applications, including offline/online metrics and regression suites.
Familiar with fine-tuning small/open models and applying recent research in agentic and episodic memory architectures.