





Tier-2 brand, metro Bangalore, mid-level (4–6 yrs), and generalist senior title raise applicant competition.
Core data engineering skills are transferable but Hudi/Presto/lakehouse expertise increases domain specificity.
Explicit 4–6 year requirement plus mandatory Hudi/Presto/Airflow/SQL/programming mandates strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and operate large-scale batch and near-real-time data pipelines powering product, business, growth, and ML use cases.
Develop and improve lakehouse architecture using Apache Hudi, implement query engines like Presto/Trino, and build orchestration workflows with Apache Airflow.
Drive data reliability, observability, SLA tracking, and optimize storage, compute, and query costs for Apna's data platform while mentoring engineers and defining architecture standards.
4-6 years of experience in data engineering, preferably at scale.
Strong hands-on experience with Apache Airflow or similar orchestration systems.
Proficient with Presto/Trino query engines and Apache Hudi concepts including copy-on-write vs merge-on-read, upserts/deletes, compaction, clustering, schema evolution, and partitioning.
Strong SQL and programming skills in Python, Java, or Scala with ability to design reliable ETL/ELT pipelines and debug complex data issues.
Experienced in large-scale distributed data processing and storage systems with a solid understanding of multiple data architectures (data warehouse, lakehouse, lambda, kappa, medallion, event-driven).
Capable of balancing trade-offs between data freshness, cost, reliability, latency, and complexity with a strong production ownership and debugging mindset.
Skilled at collaborating across product, analytics, ML, and backend engineering teams to deliver scalable platform solutions and influence engineering standards.