





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Niche ML-data specialization and funded startup brand produce moderate candidate competition in metro areas.
Core data engineering transferable, but Indic-language and multimodal audio/video expertise increases domain-specific fit requirements.
Role requires many specific technical skills and domain expertise while lacking explicit years, so filtering is moderately strict.
Build and operate scalable data ingestion and cleaning pipelines for text, audio, and video, including deduplication, quality filtering, and synthetic data generation.
Manage and version evaluation datasets to ensure accurate benchmarking and untainted data provenance, including multimodal data processing.
Design annotation guidelines, audit annotation quality, and maintain data provenance to enable reproducible and transparent dataset management.
Strong Python programming skills with experience handling datasets larger than memory, using databases and chunked/parallel processing.
Practical data engineering skills: streaming, checkpointing, handling malformed inputs, and familiarity with ML data formats (Parquet, WebDataset, HuggingFace Datasets) and object storage (S3/GCS/R2).
Work Experience Required: Not explicitly mentioned in the JD.
Genuine concern for data quality and enough ML understanding to appreciate impact of data decisions on model performance.
Experience curating training data specifically for large language models, speech, or image/video models.
Familiarity with embeddings and vector search tools (FAISS, Milvus) for semantic deduplication and retrieval.
Native or near-native fluency in one or more Indian languages, enabling quality judgment of Indic-language data, plus experience with audio/video processing or synthetic data generation is highly valued.