





Mid-level data engineer in metro locations with broad cloud and AI requirements increases competition.
Core data engineering skills transfer across industries, but Generative AI and vector DB specialization reduces portability.
Explicit 4–8 years plus mandatory Vector DBs, AWS, Airflow, Python and SQL enforce strict candidate filtering.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design and build scalable data ingestion and transformation pipelines optimized for Generative AI applications, including LLMs, RAG, and Voice AI.
Develop and manage Vector Database architectures focusing on indexing, storage optimization, metadata management, and retrieval performance of enterprise data.
Collaborate with AI Engineers, Data Scientists, and ML Engineers to support high-performance semantic search, embedding generation, and retrieval systems on AWS cloud infrastructure.
4–8 years of experience in Data Engineering with hands-on experience building scalable cloud-native data platforms.
Strong proficiency in Python and advanced SQL with experience in production-grade Apache Airflow pipeline development.
Experience with Vector Databases such as Pinecone, Milvus, Qdrant, Chroma, or Weaviate and managing large-scale enterprise SQL environments.
Strong working knowledge of AWS services including S3, Glue, EMR, Lambda, Athena, IAM, and CloudWatch; experience with Generative AI architectures and support for ML pipelines.
Experienced in building and optimizing data pipelines specifically for Generative AI and vector search workloads, signaling strong domain expertise.
Comfortable working cross-functionally with AI/ML teams and engineering teams to operationalize AI applications at scale in cloud environments.
Skilled in both deep technical implementation (Python, SQL, Cloud, Airflow) and applied knowledge of semantic search, embeddings, and RAG architectures to enhance AI platform capabilities.