





Senior, specialized multimodal role reduces applicant density despite metro location and recognizable company brand.
Highly domain-specific multimodal vision and deep learning expertise limits cross-industry transferability.
Many mandatory advanced ML, CV, distributed training and tooling requirements create strict shortlisting filters.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design, training, and optimization of large-scale vision foundation and multimodal AI models across image and video data.
Develop and deploy scalable AI solutions including Vision Transformers, Vision-Language Models, and Retrieval-Augmented Generation pipelines.
Manage large-scale visual datasets and distributed training workflows to translate research into production-ready systems.
Master’s or PhD in Computer Science, AI, Machine Learning, or related field.
Extensive experience in deep learning, computer vision, and multimodal AI systems.
Strong programming skills in Python and proficiency with PyTorch.
Work Experience Required: Not explicitly mentioned in the JD.
Expertise in foundation models such as SAM, DINOv3, CLIP, BLIP, LLaVA, or diffusion-based vision models.
Experience building semantic retrieval systems using embeddings and vector databases like FAISS, Milvus, Pinecone, or Weaviate.
Capable of leading innovation at the intersection of AI research and scalable production systems for large-scale image/video understanding.