





Mid-level generalist Data Engineer role with common stack at a known health-tech firm increases applicant competition.
Core data engineering skills are transferable, but healthcare/regulatory experience raises domain specificity.
Explicit 5+ years plus mandatory Spark, Trino/Iceberg, Airflow, and language requirements make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and operate Spark ingestion jobs landing raw healthcare data into Apache Iceberg tables with robust schema validation and bad-record quarantine.
Develop and maintain Trino SQL transform pipelines for data validation, business-rule transformations, deduplication, and aggregate table builds.
Automate Iceberg table maintenance tasks and optimize query/pipeline performance including partitioning strategy and resource usage.
Bachelor's or Master's degree in Computer Science or a related field.
Minimum 5 years of hands-on data engineering experience building scalable production data pipelines.
Proficient in SQL with strong experience in Apache Spark and Trino (or equivalent distributed SQL engines).
Experience with Apache Iceberg or similar table formats (Delta Lake, Hudi), S3-compatible storage, Parquet format, workflow orchestration (e.g., Airflow), and coding in Python and/or Java.
Experienced in end-to-end data pipeline development and operation in regulated healthcare or similarly complex environments.
Able to work hands-on across ingestion, transformation, and serving layers of a data lakehouse architecture.
Capable of automating data pipeline maintenance and implementing data-quality instrumentation at scale.