





Popular data engineering role with broad Spark/AWS requirements in a metro location.
Core data engineering skills (Spark, Python, SQL, cloud) are highly transferable across industries.
Specific mandatory Spark, Scala/PySpark, AWS, and onsite Pune requirement enforce strict technical screening.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and optimize Apache Spark jobs using Scala and/or Python (PySpark) to process large-scale structured and unstructured datasets.
Collaborate with internal and customer teams to build reliable data pipelines ensuring data quality, consistency, and performance.
Maintain and troubleshoot production data workflows using cloud platforms (AWS) and distributed storage systems like S3 and HDFS.
Strong hands-on experience with Apache Spark, Scala and/or Python (PySpark).
Experience working with cloud services, preferably AWS (S3, EMR, Glue, Kinesis, Firehose, Hive).
Mandatory onsite work from Customer Office in Pune location.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in performance tuning and optimization of Spark applications with strong SQL and big data processing skills.
Familiar with CI/CD pipelines, streaming frameworks (Spark Streaming), and workflow orchestration tools such as Apache Airflow.
Able to collaborate effectively with cross-functional teams, owning data quality and reliability in production environments.