





Remote mid-level data role, popular title, metro location, and broad skill requirements increase applicant competition.
Core data engineering skills (PySpark, SQL, dbt, Kafka) are highly transferable across industries.
Explicit 5+ years plus many mandatory data stack skills and infra requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Build and maintain high-throughput ETL/ELT pipelines ingesting data into Lakehouse architectures.
Ensure code quality with modular, reusable code using dbt and PySpark, following CI/CD practices.
Implement monitoring systems for data SLAs, optimize SQL and Spark queries, and collaborate with BI and data science teams to prepare feature-ready datasets.
Minimum 5+ years of data engineering experience involving on-call responsibilities.
Proficiency in Python, SQL, and PySpark; experience with Databricks (Delta Lake) or Snowflake and Medallion Architecture.
Hands-on experience with dbt, Apache Airflow or Prefect, and familiarity with streaming tools like Kafka, Kinesis, or Spark Streaming.
Experience with Infrastructure as Code (Terraform or CloudFormation), DuckDB, Kubernetes, and query optimization in cloud environments.
Experienced in building resilient, self-healing data pipelines with strong operational ownership including data reliability and observability.
Skilled in balancing heavy compute (PySpark/Databricks) with lightweight local compute (DuckDB/Python) for cost and latency optimization.
Comfortable managing complex orchestration and hybrid execution patterns in cloud-first data architectures with direct S3/Azure Blob file querying.