





Remote mid-level role in Bangalore with generalist data skills attracts strong applicant competition.
Core cloud and data ingestion skills are transferable, though audio/ML dataset experience adds domain specificity.
Mandatory 5+ years and required cloud, IaC, Docker, and scripting skills increase screening rigidity.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own data collection for AI model training, focusing on sourcing new audio data and integrating it into ingestion pipelines.
Operate and extend cloud infrastructure on GCP using Terraform for large-scale data ingestion.
Collaborate with AI Scientists and leadership to optimize cost, throughput, and quality of datasets and shape the AI dataset roadmap.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency in bash/Python scripting on Linux environments.
Professional experience with Docker, Infrastructure-as-Code, and at least one major Cloud Provider (GCP preferred).
Experienced in handling large-scale data ingestion and processing workflows, ideally with web crawlers.
Capable of balancing multiple tasks and shifting priorities in a fast-growing environment.
Collaborates effectively across engineering, AI research, and leadership to drive data infrastructure strategy.