





Popular Data Engineer role, metro Bangalore, and mid-level expectations increase candidate competition.
GCP and PySpark specialization is transferable across industries but requires cloud-specific expertise.
Multiple mandatory skills (GCP services, PySpark, SQL) create moderate screening strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, develop, and maintain large-scale batch and real-time data pipelines on Google Cloud Platform.
Optimize Spark and PySpark jobs for efficient large-scale data processing.
Leverage GCP services like BigQuery, Dataproc, Dataflow, Pub/Sub, Cloud Storage, and Cloud Composer to manage data workflows.
Strong proficiency in Python programming with hands-on expertise in PySpark.
Experience working with Google Cloud Platform services including BigQuery, Dataproc, Dataflow, Pub/Sub, Cloud Storage, and Cloud Composer (Airflow).
Strong SQL skills with experience in relational databases and knowledge of data lake and data warehouse concepts.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in building and optimizing both batch and real-time data pipelines in cloud environments, specifically GCP.
Skilled in distributed computing and Spark optimization techniques to handle large-scale data processing.
Familiar with version control (Git), CI/CD pipelines, and Agile development practices indicating operational maturity.