





Specialized multimodal ML role at a known firm with metro location, moderate brand increases applicant competition.
Highly domain-specific vision and multimodal expertise limits cross-industry transferability.
Requires advanced degree plus specific multimodal, CV, production ML and infra skills, so filtering will be stringent.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Lead design, development, and deployment of large-scale Vision Foundation Models and multimodal AI systems for image and video understanding.
Develop and optimize Vision-Language Models and Retrieval-Augmented Generation (RAG) pipelines using state-of-the-art architectures and vector databases.
Manage large-scale datasets and distributed training workflows to translate cutting-edge research into scalable production solutions.
Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, or related field.
Extensive experience in deep learning, computer vision, and multimodal AI systems.
Proficient in Python and PyTorch with hands-on experience in foundation models like SAM, DINOv3, CLIP, BLIP/BLIP-2, LLaVA, or diffusion-based vision models.
Work Experience Required: Not explicitly mentioned in the JD.
Expertise in Vision Transformers, self-supervised learning, and scalable ML infrastructure such as GPU acceleration and distributed environments.
Experience building semantic retrieval systems with embeddings and vector databases such as FAISS, Milvus, Pinecone, or Weaviate.
Ability to bridge research and production by developing robust, large-scale multimodal AI systems and contributing to publications or open-source projects.