Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Fine-tune and adapt open-source speech models for proprietary call audio, ensuring operationally relevant model quality such as entity and numeric accuracy and latency.
Build and manage training and evaluation pipelines for multilingual and code-mixed speech datasets, including large-scale data curation techniques like pseudo-labeling and speech enhancement.
Evaluate and select speech model architectures based on evidence, and deploy production models in collaboration with the platform team.
Minimum Requirements
3-4 years of ML work experience, including at least 18 months focused on speech or audio.
Strong proficiency in Python and PyTorch, with a track record of implementing research papers.
Hands-on experience fine-tuning at least one production speech model.
Solid understanding of speech fundamentals including mel-spectrograms, acoustic models, vocoders, encoder-decoder and transducer architectures, and evaluation methodologies.
Ideal Candidate Profile
Experienced working with messy, real-world audio rather than just clean benchmark datasets, especially in multilingual or code-mixed contexts.
Able to independently evaluate candidate architectures critically and make data-driven decisions to ship production models.
Familiar with advanced speech system components such as self-supervised encoders, neural audio codecs, and language-model-based generation techniques.
