





Mid-level, metro Data Engineer title with generalist stack increases competitive applicant density.
Core data engineering skills are transferable, though healthcare/regulatory experience moderately biases fit.
Explicit 5+ years plus mandatory Spark, Trino/Iceberg, Airflow, and CI/CD skills make shortlisting highly stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and operate large-scale data ingestion pipelines using Apache Spark to land raw healthcare data into Apache Iceberg tables with schema handling and error quarantine.
Develop, maintain, and optimize Trino SQL transform pipelines for layered data validation, business rules, deduplication, and aggregation supporting analytics and applications.
Automate table maintenance tasks and tune query and pipeline performance, including compaction, snapshot expiry, resource management, and data quality instrumentation.
Bachelor's or Master's degree in Computer Science or related technical field.
5+ years of experience in data engineering building production-scale data pipelines.
Strong skills in SQL and hands-on experience with Apache Spark for batch processing.
Experience with Trino/Presto or similar distributed SQL engines and open table formats such as Apache Iceberg (preferred), Delta Lake, or Hudi.
Experienced in end-to-end data pipeline development and operation in healthcare data environments, including handling HL7, CCDA, or claims data formats.
Proficient in performance tuning and pipeline automation in distributed storage environments using S3-compatible storage and Parquet file formats.
Skilled in workflow orchestration (Airflow or equivalent) and familiar with CI/CD practices for data pipelines, as well as programming in Python and/or Java.