Data Engineer III - Databricks, Pyspark, Python, AWS
JPMorgan Chase & Co.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessStrong Tier-1 brand, mid-level data engineer title, metro location, and broad Databricks/PySpark requirements.
Core PySpark, SQL, and AWS skills are broadly transferable across industries.
Explicit 3+ years and mandatory Databricks, PySpark, SQL, and AWS requirements enforce strict filters.
Job Description
Structured overview of role & requirementsAbout This Role
Design, develop, and maintain big data pipelines (batch and streaming) using PySpark/Spark and Databricks, focused on high-volume datasets with reliability, quality, and performance.
Lead data modelling and solution design for data products, including defining architecture, data flows, and transformation patterns for scalable, maintainable data platforms.
Perform advanced debugging and optimization of Spark workloads, develop and optimize SQL transformations, and apply data warehousing concepts using AWS services such as S3.
Minimum Requirements
3+ years applied experience in data engineering/big data engineering with formal training or certification in data engineering concepts.
Expert-level SQL skills including complex joins, window functions, CTEs, query optimization and analytical problem solving at scale.
Strong hands-on experience with Python coding in production and deep expertise in Apache Spark and PySpark including performance tuning and distributed processing fundamentals.
Experience with Databricks, batch and streaming data processing solutions, AWS S3 and other AWS data services, plus familiarity with GitHub/Bitbucket version control workflows.
Ideal Candidate Profile
Experienced in large-scale, complex data pipeline development using Spark and Databricks, capable of leading data modelling and architecture design for enterprise products.
Able to independently troubleshoot and optimize distributed Spark systems and heavy SQL workloads, with strong coding discipline and modular design skills.
Demonstrates understanding of data security and validation through use of AI-assisted tools for pipeline design, ensuring compliance and quality within financial services or similarly regulated environments.
