





Metro location and popular data-engineer title increase candidate density despite life-sciences niche.
Requires life-sciences commercial data experience, increasing domain specificity and reducing cross-industry transferability.
Explicit 6–9 years requirement plus mandatory PySpark and Airflow and domain experience raises filtering strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain scalable batch and near real-time data pipelines using PySpark, Python, SQL on cloud-native platforms (e.g., Databricks, Snowflake).
Design, implement, and operationalize Airflow DAGs with dependencies, scheduling, and monitoring for fault-tolerant data workflows with embedded data quality checks.
Ingest and harmonize life sciences commercial datasets for use cases like targeting, incentive compensation, patient journey, and market access while collaborating with cross-functional teams to ensure data usability and documentation.
6–9 years of experience building production-grade data engineering solutions with at least 3+ years hands-on PySpark and Airflow experience.
Strong skills in PySpark ETL/ELT with large datasets, optimization, and debugging performance issues.
Experience with life sciences commercial datasets such as claims, prescription, sales, rosters, call activity, payer/plan/formulary data.
Proficiency in SQL tuning and cloud data warehouses (Snowflake, Redshift, BigQuery, or Databricks SQL).
Deep knowledge of life sciences commercial data domains and related analytics use cases for pharma field force effectiveness, patient journey, and market access.
Comfortable translating architecture/design patterns into reusable, well-structured workflows and components in a modular, CI/CD environment using Git and Jenkins or similar.
Experience working in Agile environments, owning features end-to-end from design to production support, and collaborating with analytics, data science, and product teams.