





Strong employer brand, metro location, and mid-level technical data role create high competition.
Databricks, PySpark and Hadoop skills are highly transferable across industries.
Explicit Databricks/pySpark, Hadoop, and lead-level experience requirements enforce strict technical filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and develop high-performance data pipelines using Databricks (pySpark) and Hadoop ecosystem tools.
Own performance tuning, resource optimization, and cluster efficiency on Databricks platforms.
Lead data architecture decisions including storage, security, governance, and support production stability including troubleshooting and runbook maintenance.
Bachelor of Engineering in Computer Science, Information Technology, or equivalent.
Strong hands-on experience with Databricks (pySpark) and Hadoop ecosystem (Hive, HDFS, YARN, Oozie, Starburst).
Proven experience designing and maintaining complex Big Data pipelines with Python and modern SDLC practices (Git, CI/CD).
Work Experience Required: Preferred 6+ years in data engineering or large-scale distributed systems.
Experienced in large-scale distributed systems and complex data engineering environments.
Skilled in performance tuning and optimization of Spark SQL and Hive.
Comfortable working independently and collaboratively in multicultural, Agile environments with cloud and advanced data governance knowledge.