





Strong Tier-1 brand attracts applicants but niche speech AI and CUDA expertise limits the qualified pool.
Specialized speech AI, low-latency inference and CUDA skills reduce cross-industry transferability.
Explicit 6+ years, C++/CUDA expertise and speech AI experience make hiring filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own development and optimization of GPU-accelerated Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Audio Language Models (ALM), and Speech-to-Speech (S2S) systems in production.
Collaborate with model researchers to transition speech AI models from research to production readiness and build C++/Python backend implementations leveraging CUDA for GPU acceleration.
Provide advanced technical guidance to enterprise and developer customers for integration, deployment, troubleshooting, and performance optimization of speech technology solutions.
6+ years of professional experience in system software engineering or related roles in speech technologies.
Master's or BE/BTech degree in Computer Science, Computer Architecture, or related field.
Proficiency in C++ and Python programming, including debugging and performance analysis.
Experience with inference pipelines for large language models, speech recognition, and speech synthesis, and knowledge of modern model architectures like Transformers, CNNs, RNNs.
Strong expertise in real-time streaming audio processing and low-latency inference architectures for speech AI systems.
Capable of independently defining project goals and managing end-to-end development in a dynamic matrix environment.
Experience with cloud service deployment (HTTP REST, gRPC, Websockets) and hands-on debugging across multiple software layers including kernels, containers, and storage systems.