





Popular generalist Data Engineer title and metro Bangalore location increase applicant competition.
Core data engineering skills are transferable across industries but GCP/PySpark specificity increases domain sensitivity.
Mandatory GCP, PySpark, ETL pipeline expertise enforces technical filtering across applicants.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, develop, and maintain scalable ETL/ELT data pipelines using PySpark and Python on Google Cloud Platform.
Manage end-to-end data processing solutions leveraging GCP services including BigQuery, Dataproc, Dataflow, Cloud Storage, Pub/Sub, and Cloud Composer.
Optimize PySpark jobs for performance and cost, ensure data quality and compliance, and troubleshoot production data issues.
Strong programming experience in Python and PySpark for large-scale data processing.
Proven experience with multiple Google Cloud Platform services: BigQuery, Dataproc, Dataflow, Cloud Storage, Pub/Sub, Cloud Composer (Airflow).
Strong SQL skills and experience with relational databases; experience building batch and real-time data pipelines.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in optimizing distributed computing and Spark jobs for performance and scalability within GCP environments.
Familiar with CI/CD pipelines, Git, Agile workflows, and code reviews in data engineering contexts.
Capable of collaborating effectively with Data Scientists, Analysts, and Application teams to deliver high-quality, compliant data solutions.