





Metro role and known startup brand, but specialized speech/agent research focus limits applicant pool.
Specialized speech, agent evaluation, and real-time systems expertise reduces cross-industry transferability.
Strong preference for PhD, publications, and research experience creates stringent screening filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop evaluation frameworks tailored to voice agents focusing on audio-native metrics, adversarial conversational datasets, and LLM-based rubrics for task completion and error recovery.
Design and implement end-to-end observability systems correlating multi-modal data (audio packets, STT, LLM traces, tool calls, TTS) linked by conversation ID to detect cascade failures.
Build self-improvement systems by mining production failure traces, generating targeted fine-tuning data, and validating fixes using adversarial replay.
Proficiency in Python programming.
Familiarity with at least one: speech models (e.g., Whisper, Conformer), LLM tool-use/agent frameworks, or observability stacks (e.g., OpenTelemetry).
Strong research background indicated by publications in speech, dialogue systems, or human-AI interaction venues (Interspeech, ACL, NeurIPS, EMNLP) or equivalent experience.
Current PhD students in ML, NLP, or speech preferred; exceptional MS students or research engineers with publications also considered.
Demonstrated research capability with relevant publications in speech/dialogue/human-AI interaction conferences.
Experience or interest in bridging cutting-edge research with practical product shipping.
Background or familiarity with real-time systems, streaming pipelines, or telephony environments is advantageous but not mandatory.