





Remote, mid-level generalist data engineering role in Bangalore with broad skill requirements increases applicant competition.
Core data infrastructure and cloud skills are transferable, though audio dataset experience adds moderate domain specificity.
Explicit 5+ years, required cloud/IaC/tools and domain experience imply high shortlisting strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own end-to-end data acquisition and ingestion pipeline operations on GCP for audio data supporting model training.
Develop and extend cloud infrastructure using Terraform and Docker to enable scalable, low-cost data workflows at petabyte scale.
Collaborate with AI scientists and leadership to optimize cost, throughput, and quality of datasets powering next-generation AI models and products.
Bachelor's, Master's, or PhD in Computer Science or related field.
5+ years of software development industry experience.
Proficiency with bash/Python scripting in Linux environments; experience with Docker, Infrastructure-as-Code, and at least one major Cloud Provider (GCP preferred).
Work Experience Required: 5+ years in software development.
Experienced in building and operating large-scale, cloud-based data ingestion pipelines particularly on GCP.
Demonstrated ability to integrate infrastructure engineering with AI research needs to deliver cost-effective, high-quality datasets.
Track record of collaborating cross-functionally with scientists and leadership to shape data strategies for AI model training.