





Tier-1 employer, metro location, mid-level generalist data role with broad stack increases candidate competition.
Data platform and embedding engineering skills are transferable, though financial-data experience is preferable.
Explicit 6+ years requirement and production data engineering expectations make shortlisting relatively strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain GenAI tools and AI agents providing LLMs structured, live access to data pipelines via MCP tool servers.
Build and manage data ingestion pipelines into Apache Iceberg, including orchestration with Apache Airflow and semantic search capabilities using vector databases and embedding models.
Design microservices, backend APIs, and self-service ingestion APIs to enable partner teams' onboarding and contribute to production tool interfaces and system observability.
6+ years software engineering experience in production environments.
Proficiency in Python or similar programming language.
Familiarity with relational databases and SQL.
Interest or experience in data engineering (pipelines, batch processing, or data lake technologies), distributed systems, and AI/ML, especially LLMs, RAG, or embedding-based search.
Experienced in building scalable data infrastructure and engineering tools combining data pipelines with AI/ML capabilities in a production setting.
Skilled in working with distributed systems and modern data technologies like Apache Iceberg, Airflow, vector databases, and embedding models.
Comfortable collaborating across global teams and contributing to greenfield projects shaping engineering culture and technical direction.