





Strong employer brand but highly specialized Speech AI/GPU systems role limits general applicant pool.
Specialized speech AI, GPU/CUDA, and low-latency inference demand industry-specific expertise.
Explicit 6+ years and mandatory C++, CUDA, speech inference, and cloud deployment requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Drive development and optimization of GPU-accelerated speech AI systems including ASR, TTS, ALM, and S2S for production use.
Lead troubleshooting and performance tuning for real-time streaming and low-latency inference pipelines.
Collaborate with model researchers and develop backend C++/Python services and client SDKs to transition models into scalable production environments.
Masters or BE/BTech in Computer Science, computer architecture, or related field.
6+ years of relevant industry experience in system software or speech AI implementation.
Strong proficiency in C++ and Python programming with experience in designing, debugging, and performance optimization.
Experience in inference pipelines for speech recognition, speech synthesis, and large language models; knowledge of real-time streaming audio and low-latency system architectures.
Experienced in integrating and deploying speech AI models in production with knowledge of modern model architectures (Transformers, CNNs, RNNs).
Capable of independently managing project scope and collaborating effectively in dynamic matrix organizations to guide technical implementations.
Familiar with cloud service deployment technologies (HTTP REST, gRPC, Websockets) and has experience optimizing inference performance in GPU-accelerated environments.