





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Mid-level, remote data engineering role with broad cloud and scripting requirements attracts many applicants.
Requires specific data ingestion, cloud, and IaC experience, limiting cross-industry transferability.
Requires 5+ years plus GCP, Terraform, Python, Docker and infra experience, making filters stringent.
Own end-to-end data collection and ingestion pipeline for AI model training at petabyte scale.
Operate and extend cloud infrastructure on GCP using Terraform to ensure scalable, cost-efficient data workflows.
Collaborate with scientists and leadership to innovate data sourcing and define dataset roadmap to enhance AI models powering consumer and enterprise products.
BS/MS/PhD in Computer Science or related field.
5+ years of industry experience in software development.
Proficiency in bash/Python scripting on Linux, Docker, Infrastructure-as-Code, and experience with at least one major cloud provider (GCP preferred).
Experience with data ingestion pipelines; exposure to web crawlers and large-scale data processing is a plus but not mandatory.
Experienced in managing large-scale cloud data infrastructure with focus on cost-quality optimization.
Comfortable working cross-functionally with AI research scientists to align data capabilities with model training needs.
Adaptive to dynamic priorities in a high-growth, distributed startup environment focusing on AI and audio technology.