





Tier-1 brand, mid-level generalist data role with common skillset and likely metro hiring.
Core data-engineering skills are transferable, but domain-specific payments and Hadoop expertise increase sensitivity.
Explicit 5+ years and mandatory PySpark/Hadoop/SQL and data governance requirements.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and operate scalable data pipelines and curated datasets supporting analytics, reporting, and advanced modeling.
Own pipeline performance, reliability, data quality, governance, and security for batch and streaming workloads across big data platforms.
Collaborate with cross-functional teams to translate data requirements into governed data models and ensure production issue troubleshooting and operational excellence.
5+ years of experience in data engineering or big data analytics engineering.
Strong hands-on experience with Hadoop ecosystem (HDFS, Hive, Impala, YARN, Oozie) and production-grade big data pipelines.
Proficiency in PySpark, Python, SQL, and data orchestration tools (e.g., Apache Airflow).
Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
Experienced in building and optimizing high-performance data pipelines for large-scale, distributed data environments.
Skilled in data modeling, incremental processing patterns, and creating curated datasets for analytics and AI/ML use cases.
Operates with strong problem-solving skills, independently debugging complex distributed data issues and communicating effectively with diverse stakeholders.