





Remote, mid-level data role in a metro location with broad cloud and scripting requirements increases competition.
Data infrastructure and ingestion skills transfer easily across industries, so background sensitivity is low.
Explicit 5+ years requirement plus mandatory cloud, Terraform, Docker, and scripting skills raises shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the end-to-end data acquisition pipeline for model training, including identifying and integrating new audio data sources.
Operate and extend cloud infrastructure (GCP, Terraform) to support ingestion processes at petabyte scale and low cost.
Collaborate with AI Scientists and leadership to optimize dataset quality, cost, and throughput, and shape the dataset roadmap for next-generation products.
Bachelor's, Master's, or PhD in Computer Science or a related field.
5+ years of industry experience in software development.
Proficiency in bash/Python scripting in Linux environments, Docker, Infrastructure-as-Code, and experience with at least one major Cloud Provider (GCP preferred).
Experience with web crawlers and large-scale data processing workflows is a plus but not mandatory.
Experienced in managing large-scale data ingestion and cloud infrastructure with strong cross-functional collaboration with research/scientific teams.
Comfortable working in a fast-paced, evolving AI startup environment focused on dataset quality, cost optimization, and operational scalability.
Demonstrates ability to find scrappy solutions to data sourcing challenges and drive the AI team's dataset strategy aligned with product evolution.