





Remote, mid-level generalist data role attracts many applicants despite moderate employer brand.
Core data engineering skills are transferable, but audio/ML dataset acquisition expertise increases domain specificity.
Requires 5+ years plus cloud, IaC, Docker, and scripting, enforcing strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the end-to-end data ingestion pipeline for audio data to support AI model training, including sourcing new audio data and operating cloud infrastructure on GCP using Terraform.
Collaborate closely with AI scientists to optimize data quality, throughput, and cost at petabyte scale to improve next-generation speech models.
Contribute to strategic dataset roadmap planning alongside AI team and company leadership to enable consumer and enterprise product advancements.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of professional software development experience.
Proficiency in bash/Python scripting within Linux environments and experience with Docker, Infrastructure-as-Code principles, and at least one major Cloud Provider (preferably GCP).
Work Experience Required: 5+ years; Notice period: Not explicitly mentioned in the JD.
Experienced in managing large-scale data ingestion and processing workflows with cloud infrastructure and scripting automation.
Skilled in integrating engineering efforts with AI research priorities to improve data pipelines at scale.
Capable of collaborating across scientific and leadership teams to translate technical efforts into strategic product impact.