





Mid-level data engineering role in metro with common Spark/Python skills increases candidate competition.
Core data engineering skills transfer across industries, but federation and lakehouse expertise narrow suitability.
Explicit 5+ years plus mandatory federation, Spark, and lakehouse skills create stringent shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and implement a federated query layer using tools like Starburst or Trino to enable high-speed analytics across distributed data sources without data movement.
Develop scalable ETL/ELT pipelines with Python and Apache Spark, and manage modern Lakehouse table formats such as Delta Lake or Iceberg to ensure ACID transactions.
Optimize Spark jobs and SQL queries for performance; implement data governance including access control and data masking within the federation environment.
5+ years experience with Python and Apache Spark including performance tuning.
Hands-on experience with Data Federation tools such as Starburst Enterprise, Trino (Presto), or Dremio.
Proven experience with Lakehouse architectures using Delta Lake or Iceberg.
Expert SQL skills for complex analytic queries; experience with cloud platforms like AWS EMR, Azure Databricks, or GCP.
Strong expertise in architecting and operating federated data environments that unify multiple data sources with low latency.
Deep technical proficiency in Spark tuning and modern data lakehouse formats for reliable data storage with ACID guarantees.
Experience working in cloud environments integrating compute and storage, with a focus on high-performance analytics and governance controls.