





Mid-level generalist Big Data role, broad skills and metro location increase candidate competition.
Core data engineering skills transferable, but oilfield IoT and OSI PI timeseries add domain specificity.
Explicit 5+ years plus many mandatory technologies creates strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, develop, and optimize scalable data ingestion, transformation, and analytics pipelines primarily using Spark and Delta Lake technologies, including cluster tuning and processing infrastructure optimization.
Build and maintain analytical data models, KPI frameworks, and reusable frameworks for processing large-scale IoT and event-driven datasets, supporting operational reporting and business intelligence.
Collaborate with multiple stakeholders to define and implement KPIs and data products, develop real-time and batch processing solutions, and build large-scale data migration solutions from source systems like OSI PI and TimescaleDB.
Bachelor’s degree in computer science, information systems, or related field, or equivalent relevant experience.
Minimum 5 years of relevant work experience in big data engineering or similar roles.
Advanced proficiency with Spark (Databricks preferred), including cluster tuning, and experience with distributed systems, partitioning, and time series databases such as OSI PI and TimescaleDB.
Proficiency with automation/scripting (.NET, C#, Python, JavaScript, GoLang), AWS/cloud technologies, Linux/Windows OS, containers, CI/CD tools (GitHub Actions, Jenkins), and data visualization tools (Tableau, Spotfire, Grafana).
Experienced in designing and optimizing big data pipelines and analytics workloads in industrial or IoT domains, with a strong focus on Spark and time series data.
Capable of independently managing complex technical solutions across various platforms and collaborating with cross-functional teams including product, engineering, and operations.
Strong in automation, cloud infrastructure (AWS), containerization, and observability practices supporting reliable, scalable, and maintainable data solutions.