Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design, build, and maintain scalable batch and real-time streaming data pipelines across diverse data sources.
Architect dimensional data models and manage data lakes/lakehouses and various storage solutions ensuring performance and reliability.
Implement data governance, security compliance (GDPR, CCPA, SOC2, HIPAA), pipeline monitoring, and enable stakeholder self-service datasets.
Minimum Requirements
5+ years experience in complex data platform implementation and architecture decisions.
Proficient in Python and/or Scala/Java with strong software engineering practices (OOP, unit testing, git).
Advanced SQL skills with expertise in modern cloud data warehouses (Snowflake, BigQuery, Redshift, Databricks Lakehouse).
Experience with distributed computing frameworks (Apache Spark/PySpark, Flink, or Hadoop) and cloud platforms (AWS, GCP, or Azure).
Ideal Candidate Profile
Experienced in both batch and real-time data pipeline development and orchestration using tools like dbt and Apache Airflow.
Demonstrated ability to optimize query performance and implement automated data quality/testing frameworks.
Skilled in enforcing enterprise data governance policies and working closely with data consumers including BI and ML teams.
