





Popular mid-level data engineer role with broad required skills and mid experience range.
Data engineering skills are transferable across industries though AWS/Hadoop specifics add moderate bias.
Explicit 4–8 years plus mandatory Spark, AWS, and ETL tech stack increases filter strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Own the design, build, and maintenance of scalable, high-volume ETL/ELT data pipelines leveraging Hadoop ecosystem (HDFS, Hive, Spark, Kafka) and AWS data stack (Glue, EMR, Lambda, Step Functions, Redshift).
Implement and optimize distributed data processing solutions using PySpark, Spark SQL, and cloud serverless architectures, ensuring data governance and compliance.
Lead architecture decisions, troubleshoot Spark and cluster performance issues, and collaborate with cross-functional teams to deliver complex Agile projects and ensure smooth operations.
4 to 8 years of overall ETL experience, including Big Data with Spark and Cloudera.
Strong hands-on expertise with AWS data stack (S3, Glue, EMR, Lambda, Step Functions, Redshift) and Hadoop ecosystem components (HDFS, Hive, Spark, Kafka).
Proficiency in PySpark, Scala, SQL (HiveQL, Impala), and scripting languages such as Python and Shell scripting.
Experience with workflow orchestration tools like Airflow, Control-M, or AWS Step Functions; familiarity with CI/CD tools including GitHub.
Experienced in end-to-end serverless data architectures on AWS integrating Glue, Lambda, S3, and Redshift with good understanding of hybrid data lake architectures (S3 + HDFS).
Able to define technical specifications and architecture aligned with business goals, with strong problem-solving skills to troubleshoot Spark performance and job failures.
Comfortable working in Agile environments delivering complex data engineering projects and mentoring peers on best practices and secure coding standards.