





Remote role, mid-level experience band, and metro location increase applicant competition.
Data infrastructure skills are transferable across industries, though AI/audio dataset focus adds moderate specificity.
Explicit 5+ years and mandatory cloud/IaC/Docker/Python skills create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the end-to-end data collection and ingestion pipeline to support AI model training, including sourcing new audio datasets.
Operate and enhance the cloud infrastructure for data ingestion on Google Cloud Platform, managed via Terraform.
Partner closely with AI researchers to improve data quality, scale, and cost-efficiency, contributing to the AI dataset roadmap for future products.
Bachelor's, Master's, or PhD in Computer Science or a related field.
5+ years of industry experience in software development.
Proficiency with bash/Python scripting in Linux environments, Docker, Infrastructure-as-Code, and professional experience with at least one major cloud provider (GCP preferred).
Work Experience Required: 5+ years in software development.
Experienced in building scalable, large-scale data processing workflows and handling web crawlers or similar ingestion systems.
Comfortable operating at the intersection of infrastructure, engineering, and AI research to optimize data pipelines for performance and cost.
Able to collaborate closely with AI scientists and leadership to align data operations with strategic product goals.