





Mid-level competition due to common data-engineer role and required specialized Spark/GCP skills.
Medium because data engineering skills (Spark, Kafka, GCP) are transferable but cloud specialization matters.
High due to explicit 8+ years requirement and 4+ years recent GCP plus specific tech mandates.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, develop, and maintain ETL/ELT data pipelines for batch and real-time processing using Spark (PySpark/Scala) and streaming technologies like Kafka and Flink.
Build and optimize scalable data architectures including data lakes, data warehouses (BigQuery), and streaming platforms with performance tuning for efficiency and cost-effectiveness.
Implement data quality checks, monitoring systems to ensure data accuracy, consistency across real-time and batch data workflows.
Minimum 8 years total IT experience with at least 4 years of recent experience on Google Cloud Platform (GCP).
Strong programming skills in Python, SQL; knowledge of Scala or Java is required for Spark development.
Expertise in big data frameworks—Apache Spark (Spark SQL, DataFrames, Streaming) and streaming technologies such as Apache Kafka or Pub/Sub.
Familiarity with data warehousing solutions (BigQuery, Snowflake, Redshift) and cloud data services on GCP or Azure.
Experienced in building large-scale, real-time, and batch data processing pipelines within cloud environments, specifically GCP.
Capable of optimizing and tuning big data workflows focusing on performance and cost efficiency.
Hands-on with modern data engineering tools and frameworks, including Airflow, Databricks, Docker, and Kubernetes, indicating operational ownership and scalability mindset.