





Mid-level, popular Data Engineer title with broad AWS/PySpark requirements and metro location increases candidate competition.
Core PySpark, AWS, and ETL skills are highly transferable across industries despite optional healthcare exposure.
Explicit 3-5 years plus mandatory PySpark, AWS, Redshift and ETL pipeline skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, test, tune, and support production ETL and data pipelines using PySpark, Python, advanced SQL, and AWS data services.
Handle ingestion and transformation from multiple data sources including flat files, relational DBs, APIs, and enterprise data.
Implement CDC, incremental loads, idempotent processing, data reconciliation, automated tests, and release support for reliable pipelines.
Bachelor's degree or equivalent in Computer Science, IT, Data Engineering, or related field.
3-5 years of experience in data engineering, ETL development, SQL, AWS data platforms, or production data pipeline support.
Mandatory technical skills: PySpark, Python, advanced SQL, AWS data services (S3, Glue, Lambda, Step Functions, ECS, DynamoDB, Redshift, PostgreSQL, SQL Server, Athena).
Experience with CI/CD, GitHub workflows, automated testing, and release management for data pipelines.
Proven expertise in building and supporting scalable data pipelines with a strong focus on AWS ecosystem and PySpark-based ETL workflows.
Experienced with advanced data platform tuning and optimization (Redshift, Athena, DynamoDB) and secure data handling including PHI/PII considerations.
Comfortable working cross-functionally with analysts, architects, QA, DevOps, and senior engineers in Agile delivery environments.