Open-Qwen2VL開源多模態模型 - 支持圖像與文本輸入並生成文本內容

首頁

Open Qwen2VL

由weizhiwang開發

Open-Qwen2VL是一個多模態模型，能夠接收圖像和文本作為輸入並生成文本輸出。

圖像生成文本英語開源協議:CC #多模態圖文理解 #學術開源模型 #高效預訓練

下載量 568

發布時間 : 3/27/2025

模型概述

基於學術資源的高效計算全開放多模態大語言模型預訓練，支持圖像和文本輸入，生成文本輸出。

模型特點

多模態輸入

支持同時接收圖像和文本作為輸入，進行聯合理解與處理。

高效計算

基於學術資源進行高效計算，適合資源有限的研究環境。

全開放

模型、代碼和數據完全開放，便於研究和二次開發。

模型能力

圖像理解

文本生成

多模態推理

使用案例

圖像描述

圖像內容描述

對輸入的圖像進行詳細描述，生成自然語言文本。

生成準確、詳細的圖像描述文本。

視覺問答

基於圖像的問答

根據圖像內容回答相關問題。

提供與圖像內容相關的準確答案。

🚀 Open-Qwen2VL模型介紹

Open-Qwen2VL是一個多模態模型，它以圖像和文本作為輸入，並輸出文本。該模型能夠有效處理圖像與文本的信息融合，為多模態任務提供了強大的支持。

🚀 快速開始

Open-Qwen2VL是一個多模態模型，它接收圖像和文本作為輸入，並輸出文本。該模型的相關信息在論文 Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources 中有所描述。代碼可在 https://github.com/Victorwz/Open-Qwen2VL 獲取。

✨ 主要特性

多模態處理：能夠同時處理圖像和文本輸入，輸出文本結果。
開源可用：代碼、模型、數據和論文均已發佈。

📦 安裝指南

請首先通過以下命令安裝Open-Qwen2VL：

pip install git+https://github.com/Victorwz/Open-Qwen2VL.git#subdirectory=prismatic-vlms

💻 使用示例

基礎用法

import requests
import torch
from PIL import Image
from prismatic import load

device = torch.device("cuda") if torch.cuda.is_available() else torch.device("cpu")

# Load a pretrained VLM (either local path, or ID to auto-download from the HF Hub)
vlm = load("Open-Qwen2VL")
vlm.to(device, dtype=torch.bfloat16)

# Download an image and specify a prompt
image_url = "https://huggingface.co/adept/fuyu-8b/resolve/main/bus.png"
# image = Image.open(requests.get(image_url, stream=True).raw).convert("RGB")
image = [vlm.vision_backbone.image_transform(Image.open(requests.get(image_url, stream=True).raw).convert("RGB")).unsqueeze(0)]
user_prompt = "<image>\nDescribe the image."

# Generate!
generated_text = vlm.generate_batch(
    image,
    [user_prompt],
    do_sample=False,
    max_new_tokens=512,
    min_length=1,
)
print(generated_text[0])

圖像描述結果如下：

The image depicts a blue and orange bus parked on the side of a street. ...

📚 詳細文檔

模型信息

屬性	詳情
基礎模型	Qwen/Qwen2.5 - 1.5B - Instruct、google/siglip - so400m - patch14 - 384
數據集	weizhiwang/Open - Qwen2VL - Data、MAmmoTH - VL/MAmmoTH - VL - Instruct - 12M
語言	英文
許可證	cc
任務類型	圖像文本到文本

更新記錄

[2025年4月1日] 代碼庫、模型、數據和論文發佈。

致謝

本工作部分得到了美國國家科學基金會BioPACIFIC材料創新平臺的資助，資助編號為DMR - 1933487。

📄 許可證

本項目採用cc許可證。

引用

@article{Open-Qwen2VL,
    title={Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources},
    author={Wang, Weizhi and Tian, Yu and Yang, Linjie and Wang, Heng and Yan, Xifeng},
    journal={arXiv preprint arXiv:2504.00595},
    year={2025}
  }