





Tier-1 brand, mid-level generalist data role in Gurgaon with broad skillset increases applicant density.
Core big-data engineering skills (PySpark, SQL, Hadoop) transfer easily across industries.
Explicit 2.5–4 year requirement plus deep PySpark, Hadoop, and data platform skill mandates strict shortlisting.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and operate scalable, high-performance data pipelines on big data platforms for analytics, reporting, and advanced modeling.
Ensure data quality, governance, and compliance by automating checks, lineage documentation, and supporting privacy/security requirements.
Collaborate with cross-functional teams (Product, Data Science, Platform) to translate requirements into reusable, governed data models and ML-ready datasets, and troubleshoot production issues.
2.5-4 years of relevant experience in data engineering or big data analytics engineering.
Bachelor’s degree in Computer Science, Engineering or equivalent practical experience.
Proficiency in building production-grade pipelines using PySpark/Spark, Python, SQL, and big data ecosystems (Hadoop: HDFS, Hive, Impala, YARN, Oozie).
Experience with data orchestration tools (e.g., Apache Airflow, NiFi), and knowledge of DevOps/CI-CD (Git, testing, deployment automation).
Experienced data engineer with deep expertise in scalable big data pipeline development and optimization on Hadoop or cloud platforms.
Skilled in producing governed, high-quality datasets for analytics and ML/AI use cases with strong understanding of data modeling and incremental processing.
Collaborates effectively with product and data science teams to translate complex requirements into operational data workflows, communicating clearly with technical and business stakeholders.