





Remote mid-level data engineering role with broad cloud and ingestion requirements attracts high candidate density.
Core data engineering skills transferable across industries, though audio/ML dataset experience adds moderate domain bias.
Explicit 5+ years and mandatory cloud/IaC/tool proficiency create strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own data collection and ingestion pipelines to support AI model training operations at petabyte-scale with cost efficiency.
Operate and extend cloud infrastructure for data ingestion pipelines on GCP using Terraform.
Collaborate with AI scientists and leadership to enhance dataset quality, scale, and roadmap for next-gen consumer and enterprise products.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of software development experience in industry.
Proficiency in bash/Python scripting on Linux, Docker, Infrastructure-as-Code, and professional experience with GCP.
Experience with web crawlers and large-scale data processing workflows is a plus, but not mandatory.
Experienced in building and managing scalable data infrastructure in cloud environments, particularly GCP.
Able to integrate infrastructure and data engineering with AI research workflows to optimize cost and dataset quality.
Capable of designing and executing strategic data acquisition aligned with AI product roadmap and cross-functional collaboration.