Data Engineer III - Databricks, Pyspark, Python, AWS
JPMorgan Chase & Co.Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessTier-1 brand, mid-level generalist data role, metro hiring, and broad toolset increase competition.
Databricks, PySpark, Python and SQL skills are highly transferable across industries.
Explicit 3+ years plus mandatory Databricks, PySpark, SQL, and AWS requirements increase screening strictness.
Job Description
Structured overview of role & requirementsAbout This Role
Design, develop, and maintain big data pipelines (batch and streaming) using PySpark/Spark and Databricks, including building scalable ingestion and transformation workflows for high-volume datasets with reliability, quality, and performance.
Lead data modelling and solution design for data products, defining target architectures, data flows, and transformation patterns to support analytics-ready data layers.
Optimize complex SQL transformations and perform advanced debugging/troubleshooting of distributed Spark workloads, applying modular design, efficient coding, and CI-friendly development with version control workflows.
Minimum Requirements
3+ years of applied experience in data engineering / big data engineering with formal training or certification in data engineering concepts.
Expert-level SQL skills including complex joins, window functions, CTEs, query optimization at scale.
Strong hands-on programming experience with Python, Apache Spark/PySpark (including performance tuning), and Databricks for large-scale data processing.
Familiarity with AWS S3 and common AWS data processing services; experience with batch and streaming data processing solutions.
Ideal Candidate Profile
Experienced in building scalable, maintainable data platforms with strong data warehousing knowledge, including dimensional modeling and SCDs.
Capable of leading the end-to-end design and delivery of robust big data solutions working with cross-functional teams, emphasizing performance and quality.
Comfortable using enterprise-authorized AI tools for data engineering workflows with strong validation and security awareness.
