





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Generalist mid-level data role, metro location, and broad in-demand skills increase applicant competition.
Core data engineering skills are transferable, though Spark and LLM specialization increases domain specificity.
Explicit 5+ years and mandatory PySpark, AWS, SQL, and LLM expertise make filters strict.
Design, develop, and maintain scalable data pipelines and ETL/ELT workflows for large-scale structured and unstructured datasets.
Build and optimize data ingestion and data warehouses to support analytics, reporting, and real-time processing using cloud platforms (AWS, Azure, or GCP).
Implement data validation, transformation, quality checks, and integrate data solutions following software engineering best practices (Git, CI/CD, testing, monitoring).
5+ years of professional experience in Data Engineering.
Strong programming skills in Python with hands-on experience in PySpark and Apache Spark.
Advanced SQL skills including query optimization and experience with cloud services: AWS, Azure, or Google Cloud Platform.
Bachelor's degree in Computer Science, Information Technology, Engineering, or related technical discipline.
Experienced in building robust ETL/ELT pipelines and data warehouses with strong knowledge of data warehousing and dimensional modeling.
Familiar with modern data engineering tools and frameworks (e.g., Apache Airflow, Kafka, Databricks) and cloud-native data lake technologies.
Has expertise or hands-on experience in Generative AI and Large Language Models (LLMs) applied to data engineering contexts.