





Tier-1 brand, mid-level generic data role, Pune location, and broad skillset increase competition.
Core PySpark, Hadoop, and data engineering skills are broadly transferable across industries.
Explicit 4-year requirement plus mandatory PySpark/Hadoop/SQL skills indicate strict technical filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and test data-centric applications using PySpark and Python in Big Data environments.
Build and optimize complex data pipelines and workflows using Hadoop ecosystem tools such as Hive, Spark, Impala, and Sqoop.
Ensure data quality by integrating testing methodologies and delivering actionable insights through business intelligence tools.
Bachelor's degree in Computer Science, Information Systems, Engineering, or related field (or equivalent practical experience).
Minimum 4 years of experience in data development, database management, or software quality assurance roles.
Strong programming skills in PySpark and Python within a Big Data environment, with proficiency in complex SQL query writing and optimization.
Familiarity with big data tools (Hadoop, Hive, Spark, Impala, Sqoop), shell scripting, Autosys scheduler, data warehouse concepts, and experience with streaming data platforms.
Experienced in end-to-end data pipeline development across distributed systems using PySpark and the Hadoop ecosystem.
Possesses strong analytical and problem-solving skills with ability to independently manage risks and resolve issues in data projects.
Comfortable working in hybrid environments spanning data engineering and quality assurance with sound knowledge of ETL processes and version control (Git).