Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Develop and maintain reliable production big-data ingestion pipelines for on-premises and cloud-based data lakes using CI/CD and agile practices.
Collaborate with data source teams and domain experts to define, implement, and validate data ingestion, cleansing, and enrichment requirements ensuring usable datasets for advanced analytics.
Monitor system performance in a DevOps environment and stay updated with industry standards to enhance pipeline quality, productivity, and performance.
Minimum Requirements
Bachelor's degree in computer science, engineering, mathematics, or related technical discipline.
Proficiency in at least one modern JVM language (Java, Scala, Kotlin) and Python programming.
Entry-level experience with AWS platform services including S3, EC2, DMS, RDS, EMR, RedShift, Lambda, DynamoDB, CloudWatch, and CloudTrail.
Familiarity with big data technologies such as Apache Spark, S3, Parquet, Delta Lake, relational and polyglot persistence, CI/CD tools (Git, Jira, Terraform), and notebook environments like JupyterHub.
Ideal Candidate Profile
Comfortable working in cross-functional teams including data source teams, domain experts, and data scientists in an agile environment.
Experienced in implementing automated validation, profiling, and data enrichment with a focus on production-grade, maintainable data pipelines.
Familiar with DevOps monitoring and continuous integration/deployment processes ensuring pipeline resiliency and performance.
