Match Score
Against your primary resumeLogin to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Protocol Intelligence
Data-driven signals on your job's competitivenessStrong Tier-1 brand and metro location, but specialized AI data engineering reduces applicant pool.
Medium — core data engineering skills transfer, but enterprise payments and compliance needs increase specificity.
High — senior role with many mandatory cloud, Spark, streaming, orchestration, and AI-specific tech requirements.
Job Description
Structured overview of role & requirementsAbout This Role
Lead design, development, and optimization of scalable, cloud-native enterprise data platforms and pipelines supporting AI and Generative AI solutions.
Build and maintain robust batch and real-time data pipelines for AI model training, feature engineering, inference, and analytics with focus on high data quality, observability, and operational excellence.
Collaborate across teams to implement secure, compliant data governance and operationalize MLOps/AgenticOps capabilities enabling AI lifecycle workflows.
Minimum Requirements
Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, or related field.
Significant experience with enterprise-scale data engineering solutions for AI and machine learning workloads.
Strong proficiency in Python, SQL, Spark, distributed data processing, and cloud platforms (AWS, Azure, or GCP).
Experience with streaming (Kafka, Kinesis, Event Hubs), orchestration tools (Airflow, Azure Data Factory), data lakes/lakehouse, vector databases, and data governance.
Ideal Candidate Profile
Expert in building and optimizing data architectures for large language models (LLMs), retrieval-augmented generation (RAG), embeddings, and AI knowledge repositories.
Experienced in implementing automated data validation, monitoring, security, and compliance controls for enterprise AI data pipelines.
Skilled at collaborating with AI engineers and data scientists to deliver reusable data products and accelerate AI innovation in a large-scale, distributed cloud environment.
