emova-qwen-2-5-7b-hf オープンソース大規模モデル - 視聴覚および音声のマルチモーダル理解と生成をサポート

ホーム

Emova Qwen 2 5 7b Hf

Emova-ollmによって開発

EMOVAはエンドツーエンドの全モーダル対応大規模言語モデルで、外部モデルに依存せずに視覚、聴覚、音声機能をサポートし、マルチモーダル理解と生成を実現します。

マルチモーダル融合

Transformers

複数言語対応オープンソースライセンス:Apache-2.0 #感情音声対話 #マルチモーダル理解 #エンドツーエンド全モーダル対応

ダウンロード数 36

リリース時間 : 3/11/2025

モデル概要

EMOVAは全モーダル対応の大規模言語モデルで、テキスト、視覚、音声入力を処理し、感情制御付きのテキストと音声応答を生成できます。高度な視覚言語理解、感情音声対話、構造化データ理解を備えた音声対話能力を持っています。

モデル特徴

全モーダル性能

視覚言語と音声ベンチマークでリーダー級の結果を達成し、テキスト、視覚、音声の入出力をサポートします。

感情音声対話

意味-音響分離音声トークナイザーと軽量スタイル制御モジュールを採用し、24種類の音声スタイル制御（2話者、3音高、4感情）をサポートします。

多様な構成

3種類のパラメータ規模（3B/7B/72B）のモデル構成を提供し、異なる計算予算ニーズに対応します。

モデル能力

テキスト生成

画像分析

音声認識

音声合成

感情制御

マルチモーダル対話

使用事例

スマートアシスタント

感情音声アシスタント

スマートアシスタントとして、感情付きの音声応答を理解・生成でき、ユーザー体験を向上させます。

24種類の音声スタイル制御をサポートし、生き生きとした音声インタラクションを実現します。

視覚言語理解

画像キャプション生成

画像内容を分析し、詳細なテキスト説明を生成します。

DocVQAデータセットで94.2%の精度を達成しました。

音声認識と合成

音声テキスト変換

音声入力をテキスト出力に変換します。

LibriSpeech (clean)テストセットでWER 4.1を達成しました。

🚀 EMOVA-Qwen-2.5-7B-HF

EMOVA（EMotionally Omni-present Voice Assistant）は、外部モデルに依存せずに、見る、聞く、話すことができる革新的なエンドツーエンドのオムニモーダル大規模言語モデル（LLM）です。オムニモーダル（テキスト、ビジュアル、音声）の入力を受け取ると、EMOVAは音声デコーダーとスタイルエンコーダーを利用して、生き生きとした感情制御を伴うテキストと音声の両方の応答を生成することができます。EMOVAは、一般的なオムニモーダル理解と生成能力を備えており、高度なビジョン言語理解、感情豊かな音声対話、構造化データ理解を伴う音声対話において優れた性能を発揮します。主な利点を以下にまとめます。

✨ 主な機能

最先端のオムニモーダル性能：EMOVAは、ビジョン言語と音声の両方のベンチマークで同時に最先端の結果を達成しています。最も性能の高いモデルであるEMOVA-72Bは、GPT-4oやGemini Pro 1.5などの商用モデルを上回っています。
感情豊かな音声対話：意味論的・音響的分離型の音声トークナイザーと軽量なスタイル制御モジュールを採用して、シームレスなオムニモーダルアライメントと多様な音声スタイル制御を実現しています。EMOVAは、24種類の音声スタイル制御（2人の話者、3種類のピッチ、4種類の感情） を備えた両言語（中国語と英語） の音声対話をサポートしています。
多様な構成：3種類の構成（EMOVA-3B/7B/72B）をオープンソースで公開して、さまざまな計算リソースに対応したオムニモーダルな使用をサポートしています。モデルズーをチェックして、あなたの計算デバイスに最適なモデルを見つけてください！

📦 インストール

このREADMEには具体的なインストール手順が記載されていないため、このセクションをスキップします。

💻 使用例

基本的な使用法

from transformers import AutoModel, AutoProcessor
from PIL import Image
import torch

### Uncomment if you want to use Ascend NPUs
# import torch_npu
# from torch_npu.contrib import transfer_to_npu

# prepare models and processors
model = AutoModel.from_pretrained(
    "Emova-ollm/emova-qwen-2-5-7b-hf",
    torch_dtype=torch.bfloat16,
    attn_implementation='flash_attention_2', # OR 'sdpa' for Ascend NPUs
    low_cpu_mem_usage=True,
    trust_remote_code=True).eval().cuda()
processor = AutoProcessor.from_pretrained("Emova-ollm/emova-qwen-2-5-7b-hf", trust_remote_code=True)

# only necessary for spoken dialogue
# Note to inference with speech inputs/outputs, **emova_speech_tokenizer** is still a necessary dependency (https://huggingface.co/Emova-ollm/emova_speech_tokenizer_hf#install).
speeck_tokenizer = AutoModel.from_pretrained("Emova-ollm/emova_speech_tokenizer_hf", torch_dtype=torch.float32, trust_remote_code=True).eval().cuda()
processor.set_speech_tokenizer(speeck_tokenizer)

# Example 1: image-text
inputs = dict(
    text=[
        {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
        {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "What's shown in this image?"}]},
        {"role": "assistant", "content": [{"type": "text", "text": "This image shows a red stop sign."}]},
        {"role": "user", "content": [{"type": "text", "text": "Describe the image in more details."}]},
    ],
    images=Image.open('path/to/image')
)

# Example 2: text-audio
inputs = dict(
    text=[{"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]}],
    audios='path/to/audio'
)

# Example 3: image-text-audio
inputs = dict(
    text=[{"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]}],
    images=Image.open('path/to/image'),
    audios='path/to/audio'
)

# run processors
has_speech = 'audios' in inputs.keys()
inputs = processor(**inputs, return_tensors="pt")
inputs = inputs.to(model.device)

# prepare generation arguments
gen_kwargs = {"max_new_tokens": 4096, "do_sample": False} # add if necessary
speech_kwargs = {"speaker": "female", "output_wav_prefix": "output"} if has_speech else {}

# run generation
# for speech outputs, we will return the saved wav paths (c.f., output_wav_prefix)
with torch.no_grad():
    outputs = model.generate(**inputs, **gen_kwargs)
    outputs = outputs[:, inputs['input_ids'].shape[1]:]
    print(processor.batch_decode(outputs, skip_special_tokens=True, **speech_kwargs))

📚 ドキュメント

モデル情報

属性	详情
モデルタイプ	EMOVA-Qwen-2.5-7B-HFは、オムニモーダル大規模言語モデルで、画像、音声、テキストの入力に対応し、多様なタスクを実行できます。
学習データ	Emova-ollm/emova-alignment-7m、Emova-ollm/emova-sft-4m、Emova-ollm/emova-sft-speech-231k

パフォーマンス

ベンチマーク	EMOVA-3B	EMOVA-7B	EMOVA-72B	GPT-4o	VITA 8x7B	VITA 1.5	Baichuan-Omni
MME	2175	2317	2402	2310	2097	2311	2187
MMBench	79.2	83.0	86.4	83.4	71.8	76.6	76.2
SEED-Image	74.9	75.5	76.6	77.1	72.6	74.2	74.1
MM-Vet	57.3	59.4	64.8	-	41.6	51.1	65.4
RealWorldQA	62.6	67.5	71.0	75.4	59.0	66.8	62.6
TextVQA	77.2	78.0	81.4	-	71.8	74.9	74.3
ChartQA	81.5	84.9	88.7	85.7	76.6	79.6	79.6
DocVQA	93.5	94.2	95.9	92.8	-	-	-
InfoVQA	71.2	75.1	83.2	-	-	-	-
OCRBench	803	814	843	736	678	752	700
ScienceQA-Img	92.7	96.4	98.2	-	-	-	-
AI2D	78.6	81.7	85.8	84.6	73.1	79.3	-
MathVista	62.6	65.5	69.9	63.8	44.9	66.2	51.9
Mathverse	31.4	40.9	50.0	-	-	-	-
Librispeech (WER↓)	5.4	4.1	2.9	-	3.4	8.1	-

🔧 技術詳細

このREADMEには具体的な技術詳細が記載されていないため、このセクションをスキップします。

📄 ライセンス

このモデルは、Apache 2.0ライセンスの下で公開されています。

引用

@article{chen2024emova,
  title={Emova: Empowering language models to see, hear and speak with vivid emotions},
  author={Chen, Kai and Gou, Yunhao and Huang, Runhui and Liu, Zhili and Tan, Daxin and Xu, Jing and Wang, Chunwei and Zhu, Yi and Zeng, Yihan and Yang, Kuo and others},
  journal={arXiv preprint arXiv:2409.18042},
  year={2024}
}