





Mid-level ML role in metro with a known startup brand, though niche speech/LLM focus reduces applicant pool.
Specialized speech, LLM agent, eval and observability expertise limits cross-industry transferability.
Explicit 3-5 years, required multi-domain ML expertise and production research increases candidate filtering strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and develop evaluation infrastructure for voice agents, including audio-native metrics and adversarial datasets.
Implement observability systems that correlate audio, STT, LLM reasoning, tool calls, and TTS within conversations to identify cascade failures.
Build self-improvement systems that analyze failure patterns, generate targeted training data, validate fixes, and guard against regressions.
3 to 5 years experience in ML engineering, research engineering, or applied research.
Strong Python skills and knowledge of modern ML tooling.
Depth in at least two areas: speech/audio models, LLM agent systems, or evaluation/observability infrastructure.
Experience shipping solutions where research integrates with production systems.
Has experience translating academic research (papers from Interspeech, ACL, NeurIPS) into production-grade systems handling real traffic.
Demonstrated ability to critically evaluate benchmarks and research relevance to real-world applications.
Comfortable working in a small, integrated team combining research and production roles focusing on enterprise conversational AI.