





Tier-1 brand, mid-level data engineer role in metro with broad AWS/Spark requirements.
Core cloud data engineering skills (AWS, Spark, Python) are highly transferable across industries.
Explicit 3+ years and multiple mandatory AWS, Spark, Python, and CI/CD requirements enforce strict filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and maintain scalable ETL/ELT data pipelines and workflows on AWS using services such as Glue, Lambda, Step Functions, and data streaming tools like Kinesis or Kafka.
Manage and optimize large-scale data storage solutions including S3, Redshift, DynamoDB, and ensure data security, compliance, and governance using AWS Lake Formation, IAM, and encryption.
Monitor pipeline health and performance via CloudWatch and CloudTrail, troubleshoot inefficiencies, and collaborate closely in Agile teams with data consumers to deliver reliable data solutions.
3+ years of professional experience in data engineering or software engineering with a strong data focus and production-grade data pipeline delivery.
Proficient with AWS ecosystem (S3, Glue, Redshift, Athena, Lambda, CloudFormation, Kinesis, DynamoDB, Lake Formation) and languages Python and SQL; knowledge of big data tools like Apache Spark, Databricks, Hadoop, Kafka is required.
Bachelor's degree (Engineering/Computer Science preferred) or equivalent experience is mandatory.
Experience with DevOps practices including CI/CD pipelines (CloudFormation, Terraform), and strong understanding of data security, governance, and compliance (e.g. GDPR).
Experienced in translating complex, ambiguous business requirements into scalable, governed data platforms with emphasis on security and compliance.
Operates effectively within Agile, cross-functional teams collaborating with product, analytics, and data science stakeholders.
Holds strong practical expertise in end-to-end data pipeline architecture and optimization using AWS native tools in production environments handling very large data lakes/warehouses.