e5-small-unsupervisedオープンソーステキスト埋め込みモデル - テキスト類似度計算タスクに無料で使用可能

ホーム

E5 Small Unsupervised

intfloatによって開発

E5-smallの教師なしバージョン、弱教師ありコントラスト事前学習でテキスト埋め込みを生成、テキスト類似度計算などのタスクに適応

テキスト埋め込み

Safetensors

英語オープンソースライセンス:MIT #教師なしテキスト埋め込み #弱教師ありコントラスト学習 #英語意味検索

ダウンロード数 2,093

リリース時間 : 1/31/2023

モデル概要

このモデルはコントラスト学習に基づくテキスト埋め込みモデルで、テキストをベクトル表現に変換でき、主に文の類似度計算や情報検索タスクに使用されます

モデル特徴

教師なし事前学習

弱教師ありコントラスト学習で事前学習を実施、ラベルデータ不要

効率的埋め込み

384次元のコンパクトなテキスト埋め込み表現を生成

プレフィックス感応

'query:'と'passage:'プレフィックスで異なるテキストタイプを区別可能

モデル能力

テキストベクトル化

文類似度計算

情報検索

意味検索

使用事例

情報検索

ドキュメント検索

クエリに基づき関連ドキュメント段落を検索

BEIRベンチマークで良好な性能

意味分析

文類似度計算

2つの文間の意味的類似度を計算

🚀 E5-small-unsupervised

このモデルは、e5-smallに似ていますが、教師ありの微調整は行われていません。

Text Embeddings by Weakly-Supervised Contrastive Pre-training Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022

このモデルは12層で、埋め込みサイズは384です。

🚀 クイックスタート

このモデルは、MS - MARCOのパッセージランキングデータセットからクエリとパッセージをエンコードするために使用できます。

💻 使用例

基本的な使用法

import torch.nn.functional as F

from torch import Tensor
from transformers import AutoTokenizer, AutoModel


def average_pool(last_hidden_states: Tensor,
                 attention_mask: Tensor) -> Tensor:
    last_hidden = last_hidden_states.masked_fill(~attention_mask[..., None].bool(), 0.0)
    return last_hidden.sum(dim=1) / attention_mask.sum(dim=1)[..., None]


# Each input text should start with "query: " or "passage: ".
# For tasks other than retrieval, you can simply use the "query: " prefix.
input_texts = ['query: how much protein should a female eat',
               'query: summit define',
               "passage: As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
               "passage: Definition of summit for English Language Learners. : 1  the highest point of a mountain : the top of a mountain. : 2  the highest level. : 3  a meeting or series of meetings between the leaders of two or more governments."]

tokenizer = AutoTokenizer.from_pretrained('intfloat/e5-small-unsupervised')
model = AutoModel.from_pretrained('intfloat/e5-small-unsupervised')

# Tokenize the input texts
batch_dict = tokenizer(input_texts, max_length=512, padding=True, truncation=True, return_tensors='pt')

outputs = model(**batch_dict)
embeddings = average_pool(outputs.last_hidden_state, batch_dict['attention_mask'])

# normalize embeddings
embeddings = F.normalize(embeddings, p=2, dim=1)
scores = (embeddings[:2] @ embeddings[2:].T) * 100
print(scores.tolist())

高度な使用法

from sentence_transformers import SentenceTransformer
model = SentenceTransformer('intfloat/e5-small-unsupervised')
input_texts = [
    'query: how much protein should a female eat',
    'query: summit define',
    "passage: As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
    "passage: Definition of summit for English Language Learners. : 1  the highest point of a mountain : the top of a mountain. : 2  the highest level. : 3  a meeting or series of meetings between the leaders of two or more governments."
]
embeddings = model.encode(input_texts, normalize_embeddings=True)

パッケージのインストール

pip install sentence_transformers~=2.2.2

📚 ドキュメント

トレーニングの詳細

トレーニングの詳細については、こちらの論文を参照してください。

ベンチマーク評価

BEIRとMTEBベンチマークの評価結果を再現するには、unilm/e5を参照してください。

よくある質問

1. 入力テキストに"query: "と"passage: "の接頭辞を付ける必要がありますか？ はい、このモデルはそのようにトレーニングされているため、付けないと性能が低下します。以下はいくつかの経験則です。

オープンQAのパッセージ検索、即時情報検索などの非対称タスクでは、それぞれ"query: "と"passage: "を使用します。
意味的類似度、言い換え検索などの対称タスクでは、"query: "接頭辞を使用します。
線形プロービング分類、クラスタリングなど、埋め込みを特徴として使用する場合は、"query: "接頭辞を使用します。

2. 再現した結果がモデルカードに報告されている結果とわずかに異なるのはなぜですか？ transformersとpytorchの異なるバージョンが、無視できないがわずかな性能差を引き起こす可能性があります。

引用

もしこの論文やモデルが役に立った場合は、以下のように引用してください。

@article{wang2022text,
  title={Text Embeddings by Weakly-Supervised Contrastive Pre-training},
  author={Wang, Liang and Yang, Nan and Huang, Xiaolong and Jiao, Binxing and Yang, Linjun and Jiang, Daxin and Majumder, Rangan and Wei, Furu},
  journal={arXiv preprint arXiv:2212.03533},
  year={2022}
}