





Remote, mid-level data-engineer role in metro Mumbai with generalist cloud and software requirements increases competition.
Cloud and data engineering skills transfer widely, but audio dataset and ML focus require domain familiarity.
Explicit 5+ years plus mandatory cloud, IaC, Docker and scripting skills enforce moderately strict shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and manage data ingestion pipeline operations and infrastructure on GCP using Terraform to support model training datasets.
Identify and integrate new audio data sources to enhance dataset scale and quality at petabyte-level volumes.
Collaborate with AI scientists and leadership to optimize data quality, cost, throughput, and define the dataset roadmap for next-generation products.
BS/MS/PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency with bash/Python scripting in Linux environments, Docker, Infrastructure-as-Code, and experience with a major cloud provider (GCP preferred).
Experience with web crawlers and large-scale data processing workflows is a plus but not mandatory.
Experienced in building and scaling large data ingestion pipelines with cloud infrastructure expertise, particularly GCP and Terraform.
Skilled in collaborating cross-functionally with research scientists and leadership to drive data-related strategic decisions.
Able to balance operational execution with strategic planning for large-scale AI data acquisition and processing projects.