





Remote, mid-level generalist data engineering role with common cloud and scripting requirements increases applicant competition.
Core cloud, scripting, and data ingestion skills transfer across industries, but ML dataset experience favors AI companies.
Mandatory 5+ years plus cloud, Terraform, Docker, and infra/data experience makes shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own all aspects of audio data collection and ingestion pipeline operations to support AI model training at petabyte scale.
Operate and extend cloud infrastructure for data ingestion on GCP using Terraform and related tools.
Collaborate with AI Scientists and leadership to enhance data quality, cost-efficiency, and define dataset roadmap for next-gen consumer and enterprise products.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of industry software development experience.
Proficiency with bash/Python scripting in Linux environments and Docker.
Professional experience with at least one major Cloud Provider (preferably GCP) and Infrastructure-as-Code tools like Terraform.
Experienced in large-scale data ingestion and processing workflows, preferably with web crawlers.
Comfortable working cross-functionally with AI scientists and leadership to optimize data strategies.
Able to manage multiple priorities and work in a fully remote, distributed environment.