





Remote, mid-level generalist data role with metro hiring pool creates high competition.
Core data infrastructure skills are transferable, but audio/ML dataset focus increases domain specificity.
Explicit 5+ years plus mandatory cloud, Terraform, and Docker skills create medium strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own and operate the cloud-based audio data ingestion pipeline on GCP, managed with Terraform, including finding new audio data sources.
Collaborate with AI Scientists to optimize data cost, quality, and throughput for model training datasets at petabyte scale.
Contribute to the AI Team's dataset roadmap to support next-generation consumer and enterprise products at Speechify.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of professional software development experience.
Proficiency in bash/Python scripting on Linux, Docker, Infrastructure-as-Code, and at least one major Cloud Provider (GCP experience preferred).
Experience with web crawlers and large-scale data processing workflows is a plus but not mandatory.
Experienced software engineer skilled in building and maintaining scalable data ingestion infrastructure with cloud and IaC tools.
Able to collaborate closely with AI researchers and leadership to balance cost, quality, and scale in data pipelines.
Comfortable working in a fast-paced, distributed environment focused on large-scale AI data engineering in audio/text-to-speech domain.