





Remote, mid-level generalist data role with popular title and metro location increases candidate competition.
Core cloud and data-engineering skills transfer across industries, but AI/dataset focus adds moderate specialization.
Explicit 5+ years and mandated cloud, Terraform, Docker, and scripting make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead the acquisition and integration of diverse audio data sources into the ingestion pipeline to support model training.
Operate and extend cloud infrastructure on GCP using Terraform to manage data ingestion workflows at petabyte-scale.
Collaborate with AI scientists and leadership to optimize cost, scale, and quality of datasets powering next-generation speech-to-text AI products.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency with bash/Python scripting in Linux environments, Docker, Infrastructure-as-Code, and at least one major Cloud Provider (GCP preferred).
Experience with web crawlers and large-scale data processing workflows is a plus but not mandatory.
Experienced in managing scalable cloud infrastructure and automating data pipelines for AI model training.
Comfortable working cross-functionally with research scientists and leadership to align data strategies with product goals.
Demonstrates ability to creatively source new audio datasets and optimize cost-efficiency at scale.