





Mid-level, generalist Spark/PySpark data role in a metro with common skillset increases candidate competition.
Core data engineering skills like Spark, PySpark, and AWS are highly transferable across industries.
Explicit 5+ years requirement plus mandatory Spark, PySpark, Python, and AWS skills makes screening strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, develop, and optimize large-scale ETL/ELT pipelines and distributed data-processing solutions using Apache Spark and PySpark.
Develop scalable, reusable Spark/PySpark frameworks and optimize Spark jobs for performance, scalability, and memory utilization.
Work with AWS cloud data services (EMR, Glue, S3, Redshift) to build data pipelines supporting data lakes, warehouses, analytics, and downstream applications.
5+ years of hands-on Data Engineering experience.
Strong expertise in Apache Spark and PySpark development.
Proficient programming skills in Python.
Experience working with AWS data services including S3 and at least one of EMR or Glue.
Experienced in building and optimizing large-scale, high-performance ETL/ELT pipelines handling complex transformations on big data.
Comfortable working in a distributed computing environment and optimizing Spark workloads for production use.
Familiarity with cloud-native data engineering on AWS, including integration with multiple AWS services and handling production deployment challenges.