





Tier-1 brand, hybrid model, mid-level generalist role, and metro location increase applicant competition.
Data engineering skills transfer across industries but require specific Big Data tooling experience.
Explicit 4–7 years plus mandatory PySpark, Hadoop, and Autosys skills make filters strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and optimize large-scale data pipelines and distributed data systems using PySpark and Hadoop ecosystem components.
Develop and manage both real-time and batch data workflows ensuring high availability and low-latency data delivery.
Automate pipeline scheduling and orchestration using shell scripting and Autosys to improve operational reliability.
4-7 years of relevant experience in Big Data engineering and distributed data workflows.
Hands-on expertise in PySpark and practical knowledge of the Hadoop ecosystem (Hive, HDFS, Sqoop, Spark, Impala, Scala).
Proficiency in complex SQL queries for data extraction, validation, and analysis.
Competence in shell scripting and job scheduling with Autosys or equivalent tools.
Experienced in designing scalable data models and architectures aligned with data warehouse and dimensional modeling concepts.
Capable of independently diagnosing and resolving complex data engineering challenges within distributed systems.
Skilled at communicating technical concepts effectively to both technical and non-technical stakeholders.