ConvLLaVA-JP-1.3b-1280開源日語視覺語言模型 - 支持高分辨率圖像輸入對話

首頁

Convllava JP 1.3b 1280

由toshi456開發

ConvLLaVA-JP是一款支持高分辨率輸入的日語視覺語言模型，能夠就輸入圖像進行對話。

圖像生成文本

Transformers

日語#日語視覺問答 #高分辨率圖像理解 #多階段聯合訓練

下載量 31

發布時間 : 6/14/2024

模型概述

該模型結合了圖像編碼器和文本解碼器，支持1280x1280高分辨率輸入，能夠進行圖像描述生成和視覺問答等任務。

模型特點

高分辨率支持

支持1280x1280高分辨率圖像輸入，能夠捕捉更豐富的視覺細節

多階段訓練

採用三階段訓練策略，先訓練視覺投影器，再聯合訓練圖像編碼器和語言模型，最後進行微調

日語優化

專門針對日語進行訓練和優化，在日語視覺語言任務上表現良好

模型能力

圖像描述生成

視覺問答

圖像對話

高分辨率圖像理解

使用案例

圖像理解

圖像內容描述

對輸入圖像生成詳細的日語描述

能夠準確識別圖像中的物體及其關係

視覺問答

回答關於圖像內容的日語問題

在JA-VG-VQA-500和JA-VLM-Bench-In-the-Wild等基準測試中表現良好

人機交互

基於圖像的對話系統

與用戶就圖像內容進行自然語言對話

能夠理解複雜問題並給出相關回答

🚀 ConvLLaVA-JP模型卡片

ConvLLaVA-JP是一款視覺語言模型，能夠針對輸入的圖像進行對話交流。該模型使用特定的圖像編碼器和文本解碼器進行訓練，支持高分辨率圖像輸入。以下是關於該模型的詳細介紹。

🚀 快速開始

下載依賴

git clone https://github.com/tosiyuki/LLaVA-JP.git

推理示例

import requests
import torch
import transformers
from PIL import Image

from transformers.generation.streamers import TextStreamer
from llava.constants import DEFAULT_IMAGE_TOKEN, IMAGE_TOKEN_INDEX
from llava.conversation import conv_templates, SeparatorStyle
from llava.model.llava_gpt2 import LlavaGpt2ForCausalLM
from llava.train.dataset import tokenizer_image_token


if __name__ == "__main__":
    model_path = 'toshi456/ConvLLaVA-JP-1.3b-1280'
    device = "cuda" if torch.cuda.is_available() else "cpu"
    torch_dtype = torch.bfloat16 if device=="cuda" else torch.float32

    model = LlavaGpt2ForCausalLM.from_pretrained(
        model_path, 
        low_cpu_mem_usage=True,
        use_safetensors=True,
        torch_dtype=torch_dtype,
        device_map=device,
    )
    tokenizer = transformers.AutoTokenizer.from_pretrained(
        model_path,
        model_max_length=1532,
        padding_side="right",
        use_fast=False,
    )
    model.eval()

    conv_mode = "v1"
    conv = conv_templates[conv_mode].copy()

    # image pre-process
    image_url = "https://huggingface.co/rinna/bilingual-gpt-neox-4b-minigpt4/resolve/main/sample.jpg"
    image = Image.open(requests.get(image_url, stream=True).raw).convert('RGB')
    
    if device == "cuda":
        image_tensor = model.get_model().vision_tower.image_processor(image).unsqueeze(0).half().cuda().to(torch_dtype)
    else:
        image_tensor = model.get_model().vision_tower.image_processor(image).unsqueeze(0).to(torch_dtype)

    # create prompt
    # ユーザー: <image>\n{prompt}
    prompt = "貓の隣には何がありますか？"
    inp = DEFAULT_IMAGE_TOKEN + '\n' + prompt
    conv.append_message(conv.roles[0], inp)
    conv.append_message(conv.roles[1], None)
    prompt = conv.get_prompt()

    input_ids = tokenizer_image_token(
        prompt, 
        tokenizer, 
        IMAGE_TOKEN_INDEX, 
        return_tensors='pt'
    ).unsqueeze(0)
    if device == "cuda":
        input_ids = input_ids.to(device)

    input_ids = input_ids[:, :-1] # </sep>がinputの最後に入るので削除する
    stop_str = conv.sep if conv.sep_style != SeparatorStyle.TWO else conv.sep2
    keywords = [stop_str]
    streamer = TextStreamer(tokenizer, skip_prompt=True, timeout=20.0)

    # predict
    with torch.inference_mode():
        output_id = model.generate(
            inputs=input_ids,
            images=image_tensor,
            do_sample=False,
            temperature=1.0,
            top_p=1.0,
            max_new_tokens=256,
            streamer=streamer,
            use_cache=True,
        )
    """貓の隣にはノートパソコンがあります。"""

✨ 主要特性

模型類型：ConvLLaVA-JP是一個視覺語言模型，可以對輸入圖像進行對話。該模型使用laion/CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft作為圖像編碼器，llm-jp/llm-jp-1.3b-v1.0作為文本解碼器，支持1280 x 1280的高分辨率輸入。
訓練過程：
- 第一階段：使用LLaVA-Pretrain-JA對Vision Projector和Stage 5進行初始訓練。
- 第二階段：使用LLaVA-Pretrain-JA對Image Encoder、Vision Projector、Stage 5和LLM進行訓練。
- 第三階段：使用LLaVA-v1.5-Instruct-620K-JA對Vision Projector和LLM進行微調。
模型對比：與其他視覺語言模型在多個基準測試中的表現對比如下： | 模型 | JA-VG-VQA-500
(ROUGE-L) | JA-VLM-Bench-In-the-Wild
(ROUGE-L) | Heron-Bench(Detail) | Heron-Bench(Conv) | Heron-Bench(Complex) | Heron-Bench(Average) | | ---- | ---- | ---- | ---- | ---- | ---- | ---- | | Japanese Stable VLM | - | 40.50 | 25.15 | 51.23 | 37.84 | 38.07 | | EvoVLM-JP-v1-7B | 19.70 | 51.25 | 50.31 | 44.42 | 40.47 | 45.07 | | Heron BLIP Japanese StableLM Base 7B llava-620k | 14.51 | 33.26 | 49.09 | 41.51 | 45.72 | 45.44 | | Heron GIT Japanese StableLM Base 7B | 15.18 | 37.82 | 42.77 | 54.20 | 43.53 | 46.83 | | llava-jp-1.3b-v1.0-620k | 12.69 | 44.58 | 51.21 | 41.05 | 45.95 | 44.84 | | llava-jp-1.3b-v1.1 | 13.33 | 44.40 | 50.00 | 51.83 | 48.98 | 50.39 | | ConvLLaVA-JP-1.3b-768 | 12.05 | 42.80 | 44.24 | 40.00 | 48.16 | 44.96 | | ConvLLaVA-JP-1.3b-1280 | 11.88 | 43.64 | 38.95 | 44.79 | 41.24 | 42.31 |

📦 安裝指南

下載項目依賴：

git clone https://github.com/tosiyuki/LLaVA-JP.git

📚 詳細文檔

更多信息請參考：https://github.com/tosiyuki/LLaVA-JP/tree/main

📄 許可證

本模型使用的許可證為cc-by-nc-4.0。

精選推薦AI模型

Llama 3 Typhoon V1.5x 8b Instruct

專為泰語設計的80億參數指令模型，性能媲美GPT-3.5-turbo，優化了應用場景、檢索增強生成、受限生成和推理任務

Cadet-Tiny是一個基於SODA數據集訓練的超小型對話模型，專為邊緣設備推理設計，體積僅為Cosmo-3B模型的2%左右。

Roberta Base Chinese Extractive Qa

基於RoBERTa架構的中文抽取式問答模型，適用於從給定文本中提取答案的任務。

智啟未來，您的人工智能解決方案智庫