





Remote role, mid-level generalist title, metro location and broad skills drive high competition.
Requires ML dataset and cloud ingestion expertise, making backgrounds outside data/ML less transferable.
Explicit 5+ years requirement plus mandatory cloud, IaC, and scripting skills tighten shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Identify and acquire new audio data sources to feed into the data ingestion pipeline.
Manage, operate, and extend the cloud infrastructure (GCP, Terraform) supporting data ingestion.
Collaborate with AI scientists and leadership to optimize dataset quality, scale, cost, and to plan dataset roadmap for next-gen models and products.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency in bash/Python scripting in Linux environments; experience with Docker, Infrastructure-as-Code, and at least one major cloud provider (GCP preferred).
Work Experience Required: 5+ years relevant software development experience.
Experienced in scalable data acquisition and processing, ideally with knowledge of web crawlers or large data workflows.
Comfortable working at the intersection of infrastructure engineering and AI data needs with a focus on cost, quality, and throughput tradeoffs.
Capable of independently managing multiple complex tasks and collaborating cross-functionally with researchers and leadership.