





Remote role, mid-level generalist title, and metro location increase applicant density.
Core cloud and infra skills transfer broadly, but dataset and audio acquisition focus requires domain familiarity.
Explicit 5+ years plus mandatory cloud, IaC, Docker and scripting skills create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end data collection and ingestion pipeline to support AI model training operations.
Operate and extend cloud infrastructure on GCP using Terraform for data ingestion at petabyte scale.
Collaborate with AI scientists and leadership to optimize data quality, cost, and scale, and to define the dataset roadmap for next-gen products.
Bachelor's, Master's or PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency in bash/Python scripting on Linux, Docker, Infrastructure-as-Code, and professional experience with a major cloud provider (GCP preferred).
Experience with web crawlers and large-scale data processing workflows is a plus but not mandatory.
Experienced software engineer comfortable with cloud infrastructure, automation, and scripting for large-scale data workflows.
Able to balance cost, throughput, and quality tradeoffs in building petabyte-scale datasets.
Collaborative mindset working closely with researchers and leadership to shape AI data strategy and pipeline evolution.