





High: mid-level generalist data role with broad AWS/PySpark requirements and likely metro hiring.
Low because core PySpark, AWS, and ETL skills are highly transferable across industries.
High due to explicit 4–8 year requirement and mandatory AWS, Spark, and ETL expertise.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and maintain high volume ETL/ELT data pipelines on Hadoop and AWS ecosystems using tools like Spark, Glue, Lambda, and Redshift.
Implement and optimize distributed data processing solutions with PySpark, Spark SQL, and serverless cloud architectures ensuring scalability and reliability.
Drive best practices in code quality, data governance, orchestration (Airflow, Control-M, Step Functions), troubleshooting Spark performance, and collaborate closely with business and cross-functional teams.
4 to 8 years of overall ETL experience with hands-on expertise in Big Data technologies like Spark and Hadoop ecosystem (HDFS, Hive, YARN, Kafka).
Strong technical skills on AWS data stack: Glue, EMR, Lambda, Step Functions, Redshift, S3, Kinesis.
Proficient in Scala, PySpark, HiveQL, SQL (Hive and Impala), Python or Shell scripting, and experience with CI/CD processes using Git.
Experience with job orchestration tools such as Airflow or Control-M and knowledge of cloud security and compliance aspects.
Experienced in architecting and delivering scalable end-to-end data solutions in Agile environments with a focus on robust, secure, and maintainable code.
Capable of defining technical specifications aligned to business goals, strong analytical and problem-solving skills to optimize data workflows and troubleshoot Spark and cluster performance issues.
Familiar with hybrid data lake architectures, data modeling techniques (star/snowflake schemas), and experienced in working with cross-functional teams and stakeholders for high-quality datasets.