





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Remote role, popular Data Engineer title, mid-level (3–6 yrs), and Bangalore market create high applicant competition.
Role requires niche speech/ASR/TTS data experience and Indic language fluency, so industry transferability is low.
Many explicit must-haves (Python, audio tooling, pipelines, governance, dataset ownership) make filters strict and technical.
Own end-to-end speech and language data management for AI model training across 22+ Indian languages, including sourcing, cleaning, labeling, versioning, cataloguing, and compliance.
Build and maintain data pipelines for audio processing and text normalization; ensure data quality and annotation accuracy including running quality checks and managing annotation vendors.
Serve as the single point of contact for AI researchers and engineers for data needs, enforcing data governance, split hygiene, and consent compliance.
3-6 years of experience in data engineering, data analysis, or ML data work with real dataset ownership.
Strong Python skills (pandas or polars), scripting and automation experience at scale, plus familiarity with audio tooling (ffmpeg, sox, librosa/torchaudio).
Experience in data pipelines, object storage, dataset versioning, train/dev/test split design, and ability to manage annotation quality.
Work Experience Required: 3-6 years relevant experience as stated; Notice Period: Not explicitly mentioned in the JD.
Experience working hands-on with speech or audio data — including segmentation, alignment, transcription, or annotation operations.
Familiarity with Indian languages beyond English, sensitivity to dialects and code-switching is a plus.
Comfortable owning data governance and compliance in regulated environments, with high ownership in fast-moving AI/ML product settings.