





Mid-level, popular Data Engineer role with common AWS/PySpark skills increases applicant competition.
Core PySpark and AWS data engineering skills are easily transferable across industries.
Requires 5+ years and mandatory PySpark, Python, and AWS experience.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize pySpark applications using Spark Dataframes in Python to process large volumes of data.
Utilize AWS analytics (EMR, Athena, Glue), compute (Lambda, EC2), and storage (S3) services in data engineering solutions.
Manage version control using Git and implement efficient data processing workflows with big data technologies.
5+ years of relevant experience including hands-on Big Data technologies.
Hands-on experience with Python and PySpark, specifically building pySpark applications.
Experience with AWS services: Amazon EMR, Athena, Glue, Lambda, EC2, S3.
Experience with version control tools like Git.
Experienced in optimizing Spark jobs for large-scale data processing.
Familiarity with data warehousing concepts such as dimensions, facts, star and snowflake schemas.
Knowledge of columnar storage formats like Parquet and compression techniques like Snappy and Gzip.