Python Spark SQL Data Engineer with AWS
Persistent Systems LimitedMatch Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design, develop, and maintain batch and streaming data pipelines using Python, PySpark, and Apache Spark.
Ensure data quality and implement ETL processes for migrating and deploying data across various enterprise systems.
Collaborate with cross-functional teams to deliver scalable data solutions that meet business requirements.
Minimum Requirements
8 to 12 years of relevant work experience in Python Spark SQL and AWS data engineering.
Proficiency in Python, PySpark, Apache Spark, and SQL including advanced features like CTEs and window functions.
Hands-on experience with AWS services such as EMR, S3, Lambda, EC2, and Athena.
Bachelor’s degree in Computer Science, Information Technology, Engineering, or related discipline.
Ideal Candidate Profile
Strong understanding of Spark architecture, Spark DStreams, and Spark Structured Streaming indicating deep domain expertise.
Experience implementing and supporting both batch and streaming data pipelines operating at enterprise scale.
Familiarity with big data tools and frameworks like Kafka, data warehousing, data lakes, and Splunk as part of an advanced data engineering environment.
