OpenVLThinker-7B開源視覺語言推理模型 - 免費部署專解視覺數學問題

首頁

Openvlthinker 7B

由ydeng9開發

OpenVLThinker-7B 是一個專為處理多模態任務設計的視覺語言推理模型，特別針對視覺數學問題解決進行了優化。

圖像生成文本

Transformers

開源協議:Apache-2.0 #視覺數學推理 #多模態理解 #高精度視覺語言模型

下載量 594

發布時間 : 3/20/2025

模型概述

基於 Qwen2.5-VL-7B-Instruct 的視覺語言推理模型，專注於解決複雜的視覺數學問題，具備多模態理解和推理能力。

模型特點

多模態推理

能夠同時處理視覺和文本信息，進行跨模態推理

視覺數學問題解決

特別優化用於解決需要視覺理解的數學問題

高效推理

支持 flash_attention_2 實現高效推理

模型能力

圖像理解

文本生成

視覺數學問題解答

多模態推理

使用案例

教育

視覺數學題解答

幫助學生解答包含圖表和圖像的數學問題

準確理解題目並給出解答

研究

多模態推理研究

用於視覺語言推理相關研究

🚀 OpenVLThinker-7B

OpenVLThinker-7B 是一個視覺語言推理模型，旨在處理多模態任務，尤其針對視覺數學問題解決進行了調優。

🚀 快速開始

模型信息

屬性	詳情
基礎模型	Qwen/Qwen2.5-VL-7B-Instruct
許可證	apache-2.0
庫名稱	transformers
任務類型	圖像文本到文本

💻 使用示例

基礎用法

from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
import torch
from qwen_vl_utils import process_vision_info
import requests
from PIL import Image

# 1. Define model and processor names
model_name = "ydeng9/OpenVLThinker-7B"
processor_name = "Qwen/Qwen2.5-VL-7B-Instruct"

# 2. Load the OpenVLThinker-7B model and processor
device = "cuda:0" if torch.cuda.is_available() else "cpu"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
    device_map=device
)
processor = AutoProcessor.from_pretrained(processor_name)

# 3. Define a sample image URL and an instruction
image_url = "https://example.com/sample_image.jpg"  # replace with your image URL
instruction = "Example question"

# 4. Create a multimodal prompt using a chat message structure
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image_url},
            {"type": "text", "text": instruction},
        ],
    }
]

# 5. Generate a text prompt from the chat messages
text_prompt = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

# 6. Process image (and video) inputs from the messages
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text_prompt],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(device)

# 7. Generate the model's response (with specified generation parameters)
generated_ids = model.generate(
    **inputs,
    do_sample=True,
    max_new_tokens=2048,
    top_p=0.001,
    top_k=1,
    temperature=0.01,
    repetition_penalty=1.0,
)

# 8. Decode the generated tokens into human-readable text
generated_text = processor.batch_decode(
    generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False
)[0]

# 9. Print the generated response
print("Generated Response:")
print(generated_text)

引用信息

@misc{deng2025openvlthinker,
      title={OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement}, 
      author={Yihe Deng and Hritik Bansal and Fan Yin and Nanyun Peng and Wei Wang and Kai-Wei Chang},
      year={2025},
      eprint={2503.17352},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2503.17352}, 
}