





Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Tier-1 brand, mid-level data engineer title, metro location, and common 2–6 years experience drive high competition.
Core data engineering skills are broadly transferable, though biotech domain knowledge moderately increases specificity.
Explicit 2–6 years plus mandatory Databricks/PySpark and Delta Lake skills make filters highly strict.
Develop, test, maintain, and optimize scalable ETL/ELT data pipelines using Databricks, PySpark, and Python.
Ingest, transform, and process structured and semi-structured data for analytics, reporting, and AI/ML use cases.
Collaborate with cross-functional teams to deliver reliable datasets and support troubleshooting of data issues and pipeline failures.
2-6 years of work experience in relevant data engineering or related roles.
Bachelor’s degree in Computer Science, Data Engineering, Information Systems, Engineering, Mathematics, or related field.
Proven hands-on experience with Python, PySpark, Databricks (notebooks, clusters, jobs, Delta tables) for building and maintaining data pipelines.
Working knowledge of SQL, Delta Lake, data formats (CSV, JSON, Parquet), and cloud platforms (AWS, Azure, or GCP).
Experience optimizing Spark jobs and Databricks workflows for performance and cost efficiency.
Familiarity with AI/ML data preparation such as feature engineering and model input creation.
Comfortable working in Agile teams, supporting ETL workflows, and managing data quality, version control (Git), and documentation.