





Mid-level, generalist Data Engineer role in metros with broad, popular skillset increases candidate competition.
Core data engineering and LLM skills transfer across industries, though enterprise AI experience creates moderate domain bias.
Mandatory 3+ years plus specific AWS, PySpark, and LLM tool experience makes shortlisting highly strict.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Design, build, and maintain high-performance ETL/ELT data pipelines on AWS using PySpark, Python, SQL across EMR, Glue, Lambda, Step Functions, Redshift, and S3.
Develop, fine-tune, and deploy large language models (LLMs) such as GPT-4 class, BERT, T5 for enterprise applications incorporating Retrieval-Augmented Generation (RAG) using vector databases and frameworks like FAISS, Elasticsearch, LangChain, LlamaIndex.
Contribute to AI platform architecture, governance, data security, and collaboration on AI/ML model lifecycle frameworks, supporting enterprise AI programs with tools like Snowflake Cortex and Databricks Lakehouse AI.
3+ years of professional experience in AI/ML and Data Engineering.
Proficient in Python, PySpark, Scala, and advanced SQL.
Hands-on experience with AWS EMR and related cloud services for data engineering.
Work authorization or willingness to work in Pune, Maharashtra, India (location explicitly mentioned).
Experienced in deploying and operationalizing LLMs and RAG workflows in enterprise environments, indicating a strong AI engineering skillset.
Comfortable working with modern AI/cloud platforms like Snowflake Cortex and Databricks Genie for scalable AI workloads.
Capable of contributing to data platform architecture and governance, implying a strategic understanding of data security and model lifecycle management.