Data Engineer – PySpark, Hadoop, Hive & Big Data Pipeline Development
SynechronMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design, develop, optimize, and maintain scalable PySpark data pipelines and Cloudera Data Platform (CDP) ingestion and transformation workflows handling large-scale datasets.
Implement data validation, quality checks, and monitor data workflows to ensure availability, integrity, performance, and reliability of data processing jobs.
Collaborate with cross-functional teams to translate data requirements into reliable technical solutions and support deployment and operational readiness activities.
Minimum Requirements
6–10 years of professional experience in data engineering, big data engineering, ETL development, or related roles.
Strong hands-on experience with PySpark, Cloudera Data Platform (CDP), Apache Spark, Hadoop, Hive, and HDFS.
Bachelor’s degree in Computer Science, Engineering, IT, Data Engineering, or related field; equivalent relevant professional experience may be considered.
Work Experience Required: 6–10 years; Notice Period: Not explicitly mentioned in the JD.
Ideal Candidate Profile
Experienced in designing and troubleshooting scalable, production-grade data pipelines in large distributed or cloud-based data ecosystems.
Skilled in performance tuning, job execution, and operational monitoring within big data environments using PySpark and CDP.
Capable of collaborating with diverse teams to deliver data-driven solutions aligned with business and technology needs, emphasizing reliability and data quality.
