





Tier-1 brand, mid-level generalist role, and metro location increase candidate competition.
Big Data tooling is transferable but mandates specific Hadoop/PySpark experience, so moderate sensitivity.
Explicit 4–7 years plus mandatory PySpark/Hadoop/streaming skills create rigid shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and maintain scalable large-scale data pipelines and distributed data systems using PySpark and the Hadoop ecosystem.
Develop and manage real-time and batch data workflows with streaming platforms to ensure high availability and low latency.
Automate pipeline scheduling and orchestration for operational reliability and independently resolve technical risks and data issues.
4 to 7 years of relevant Big Data engineering experience.
Hands-on expertise in PySpark, Hadoop ecosystem components (Hive, HDFS, Sqoop, Spark, Impala, Scala) in production.
Proficient in complex SQL for large distributed data systems.
Experience with shell scripting and job scheduling tools like Autosys.
Experienced in building and optimizing distributed data workflows at scale within large organizations.
Strong technical problem solver with ability to independently diagnose and resolve complex data engineering challenges.
Comfortable working in hybrid on-site/off-site setting based in Pune and operating with cross-functional teams.