Nomic Embed Vision v1 Open-Source Visual Embedding Model - High Performance Empowers Multimodal Application Development

Nomic Embed Vision V1

Developed by nomic-ai

High-performance vision embedding model, sharing the same embedding space with nomic-embed-text-v1, supporting multimodal applications

Text-to-Image

Transformers

EnglishOpen Source License:Apache-2.0 #Multimodal Embedding #Zero-shot Learning #Cross-modal Retrieval

Downloads 2,032

Release Time : 5/13/2024

Model Overview

nomic-embed-vision-v1 is a vision embedding model capable of converting images into embedding vectors and aligning them with text embedding space for multimodal retrieval and analysis.

Model Features

Multimodal Support

Shares the same embedding space with nomic-embed-text-v1, enabling joint retrieval and analysis of text and images.

High Performance

Outperforms models like OpenAI CLIP and Jina CLIP in benchmarks such as Imagenet zero-shot, Datacomp, and MTEB.

Easy Integration

Provides simple APIs and Python clients for quick generation of image embedding vectors.

Model Capabilities

Image feature extraction

Multimodal retrieval

Text-to-image retrieval

Image classification

Use Cases

Information Retrieval

Multimodal RAG

In Retrieval-Augmented Generation (RAG) scenarios, combines text and images for multimodal retrieval.

Improves retrieval accuracy and relevance.

Data Visualization

CC3M Dataset Visualization

Visualizes 100K samples of the CC3M dataset using Nomic Atlas maps, comparing visual and text embedding spaces.

Intuitively displays the distribution and relationships of multimodal data.

🚀 nomic-embed-vision-v1: Expanding the Latent Space

nomic-embed-vision-v1 is a high - performing vision embedding model that shares the same embedding space as nomic-embed-text-v1. All Nomic Embed Text models are now multimodal!

Name	Imagenet 0 - shot	Datacomp (Avg. 38)	MTEB
`nomic-embed-vision-v1.5`	71.0	56.8	62.28
`nomic-embed-vision-v1`	70.7	56.7	62.39
OpenAI CLIP ViT B/16	68.3	56.3	43.82
Jina CLIP v1	59.1	52.2	60.1

🚀 Quick Start

The easiest way to get started with Nomic Embed is through the Nomic Embedding API.

✨ Features

High - performing vision embedding model.
Shares the same embedding space as nomic-embed-text-v1.
All Nomic Embed Text models are multimodal.

💻 Usage Examples

Basic Usage

Generating embeddings with the nomic Python client is as easy as

from nomic import embed
import numpy as np

output = embed.image(
    images=[
        "image_path_1.jpeg",
        "image_path_2.png",
    ],
    model='nomic-embed-vision-v1',
)

print(output['usage'])
embeddings = np.array(output['embeddings'])
print(embeddings.shape)

For more information, see the API reference

Advanced Usage

Using Transformers

import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel, AutoImageProcessor
from PIL import Image
import requests

processor = AutoImageProcessor.from_pretrained("nomic-ai/nomic-embed-vision-v1")
vision_model = AutoModel.from_pretrained("nomic-ai/nomic-embed-vision-v1", trust_remote_code=True)

url = 'http://images.cocodataset.org/val2017/000000039769.jpg'
image = Image.open(requests.get(url, stream=True).raw)

inputs = processor(image, return_tensors="pt")

img_emb = vision_model(**inputs).last_hidden_state
img_embeddings = F.normalize(img_emb[:, 0], p=2, dim=1)

Multimodal Retrieval


def mean_pooling(model_output, attention_mask):
    token_embeddings = model_output[0]
    input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
    return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)

sentences = ['search_query: What are cute animals to cuddle with?', 'search_query: What do cats look like?']

tokenizer = AutoTokenizer.from_pretrained('nomic-ai/nomic-embed-text-v1')
text_model = AutoModel.from_pretrained('nomic-ai/nomic-embed-text-v1', trust_remote_code=True)
text_model.eval()

encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')

with torch.no_grad():
    model_output = text_model(**encoded_input)

text_embeddings = mean_pooling(model_output, encoded_input['attention_mask'])
text_embeddings = F.normalize(text_embeddings, p=2, dim=1)

print(torch.matmul(img_embeddings, text_embeddings.T))

📚 Documentation

Data Visualization

Click the Nomic Atlas map below to visualize a 100,000 sample CC3M comparing the Vision and Text Embedding Space!

Training Details

We align our vision embedder to the text embedding by employing a technique similar to LiT but instead lock the text embedder!

For more details, see the Nomic Embed Vision Technical Report (soon to be released!) and corresponding blog post

Training code is released in the contrastors repository

Usage Note

Remember nomic-embed-text requires prefixes and so, when using Nomic Embed in multimodal RAG scenarios (e.g. text to image retrieval), you should use the search_query: prefix.

📄 License

This project is licensed under the apache - 2.0 license.

👥 Join the Nomic Community

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご