





Metro-based, common Data Engineer role with broadly required skills increases candidate competition.
GCP and PySpark specialization moderately limits cross-industry transferability while remaining broadly applicable.
Mandatory PySpark, GCP, SQL and pipeline experience make shortlisting filters strict and technical.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, develop, and maintain scalable ETL/ELT data pipelines primarily using PySpark and Python on Google Cloud Platform.
Build and manage data ingestion and processing solutions across various GCP services including BigQuery, Dataproc, Dataflow, Pub/Sub, Cloud Storage, and Composer (Airflow).
Optimize data workflows for performance, scalability, and cost efficiency while ensuring data quality, governance, and compliance standards.
Strong programming experience in Python and PySpark for large-scale data processing.
Solid experience with Google Cloud Platform services: BigQuery, Dataproc, Dataflow, Cloud Storage, Pub/Sub, Cloud Composer (Airflow).
Proficient in SQL and relational databases with experience building batch and real-time data pipelines.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in optimizing Spark jobs and distributed computing architectures for cost and performance efficiencies.
Familiarity with cloud-native data engineering practices and working collaboratively with Data Scientists, Analysts, and Application teams.
Exposure to DevOps practices including Git, CI/CD pipelines, and Agile methodologies.