





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Highly specialized voice-LLM production skillset and senior bar reduces candidate pool.
Requires specialized voice-LLM production experience, limiting cross-industry transferability.
Strict production voice-agent, model-serving and metrics requirements make shortlisting highly selective.
Design, build, and own the end-to-end production-grade Voice AI agent pipeline, including VAD, streaming STT, LLM dialogue flow, streaming TTS, barge-in/interruption handling, and telephony integration.
Manage end-to-end latency budgets and GPU serving optimization (batching, quantization, concurrency) for real-time voice agents.
Implement and operate agentic logic for conversation management including routing, tool-calling, escalation, multi-step reasoning, and performance metrics monitoring (WER, latency percentiles, concurrency, success/failure rate).
Proven experience building and operating a production voice agent handling real user/customer calls at scale, beyond demos or prototypes.
Hands-on expertise with the full stack: VAD/turn detection, streaming STT, LLM-based dialogue orchestration, streaming TTS, barge-in handling.
Experience serving LLM/STT/TTS models on GPU infrastructure using frameworks like vLLM, Triton; no exclusive use of hosted model APIs.
Work Experience Required: Not explicitly mentioned in the JD.
Strong expertise with voice orchestration frameworks (e.g., Pipecat, LiveKit Agents) and practical problem-solving on real limitations encountered in production.
Deep understanding of latency, concurrency, and scalability tradeoffs in AI voice agent pipelines with demonstrated incident management experience.
Experience with GPU-based model serving infrastructure and managing complex agentic dialogue flows optimizing real-time performance.