Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessLog in to see why each signal reads the way it does.
Job Description
Structured overview of role & requirementsAbout This Role
Design, develop, and maintain scalable, high-performance data ingestion and ETL/ELT pipelines using Java and Apache Spark for batch and real-time streaming applications.
Optimize distributed Spark workloads through profiling, debugging, and tuning memory, partitioning, and data serialization to improve latency and resource utilization.
Collaborate with cross-functional teams and implement robust data architectures including data modeling, validation, encryption, and build CI/CD pipelines enforcing quality testing standards.
Minimum Requirements
3-6 years professional software engineering experience with 2-4 years specifically in Java and Apache Spark distributed data processing application development.
Bachelor’s or Master’s degree in Computer Science, IT, Software Engineering, or a related technical field.
Proficiency in Core Java (Java 8/11/17+), Apache Spark (RDDs, DataFrames, Spark SQL, Structured Streaming), distributed systems (HDFS, YARN), SQL and NoSQL databases, messaging platforms (Apache Kafka, RabbitMQ).
Experience in test-driven development (JUnit, Mockito) and version control with Git; build tools like Maven or Gradle.
Ideal Candidate Profile
Experienced hands-on developer skilled in deeply optimizing distributed data processing pipelines and managing cluster resource usage.
Background in building resilient data pipelines integrating various storage systems and streaming technologies in cloud or distributed environments.
Comfortable working in Agile teams executing CI/CD workflows and enforcing automated testing in Java/Spark ecosystems.
