





Mid-level data engineer in metro at known pharma with common Databricks/PySpark skills, moderate competition.
Core data engineering skills transfer across industries, but healthcare regulatory experience increases domain specificity.
Explicit 5-9 years plus mandatory Databricks, PySpark, Airflow, and AWS required.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and operate large-scale healthcare data pipelines from ingestion to conformed data products focusing on reliability, data quality, and observability.
Design, maintain, and optimize PySpark/SQL pipelines in Databricks; build and support Airflow workflows for orchestration and production operations.
Collaborate with analytics, business, and platform teams to deliver trusted data sets for healthcare use cases including sales, claims, patient, and rare diseases.
Bachelor’s degree in Computer Science, Information Technology, or a related field.
5-9 years of relevant data engineering experience.
Proficiency with Python, PySpark, SQL (including window functions, complex joins, MERGE patterns), and Databricks.
Hands-on experience with Airflow, AWS cloud data platforms, object storage, secure secret handling, and data quality engineering in regulated environments.
Experienced in building and maintaining complex ETL pipelines with a focus on operational reliability and SLA-driven delivery.
Familiar with healthcare data domains or regulated data contexts requiring strong data quality and monitoring controls.
Capable of working closely with cross-functional teams such as product owners and business analysts to translate requirements into robust engineering solutions.