





Remote and mid-level software-engineer title increase competition, but dataset/infrastructure specialization moderates it.
Role requires data infrastructure and cloud experience for ML datasets, making background moderately domain-specific.
Explicit 5+ years and required cloud/IaC/docker skills impose moderately strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the sourcing and ingestion of large-scale audio data to support AI model training operations.
Operate and extend cloud infrastructure for data ingestion pipelines on Google Cloud Platform using Terraform.
Collaborate with AI scientists and leadership to optimize data quality, cost, and throughput and shape the dataset roadmap for next-gen products.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency with bash/Python scripting in Linux environments, Docker, and Infrastructure-as-Code on major cloud platforms (experience with GCP required).
Experience with web crawlers and large-scale data processing workflows is a plus but not mandatory.
Experienced engineer comfortable working in cloud infrastructure and data ingestion at petabyte scale.
Able to collaborate cross-functionally with AI researchers to balance cost, quality, and scale tradeoffs in data pipelines.
Strategic thinker contributing to long-term dataset roadmap aligned with product and AI model needs.