OpenVLThinker-7B开源视觉语言推理模型 - 免费部署专解视觉数学问题

首页

Openvlthinker 7B

由 ydeng9 开发

OpenVLThinker-7B 是一个专为处理多模态任务设计的视觉语言推理模型，特别针对视觉数学问题解决进行了优化。

图像生成文本

Transformers

开源协议:Apache-2.0 #视觉数学推理 #多模态理解 #高精度视觉语言模型

下载量 594

发布时间 : 3/20/2025

模型简介

基于 Qwen2.5-VL-7B-Instruct 的视觉语言推理模型，专注于解决复杂的视觉数学问题，具备多模态理解和推理能力。

模型特点

多模态推理

能够同时处理视觉和文本信息，进行跨模态推理

视觉数学问题解决

特别优化用于解决需要视觉理解的数学问题

高效推理

支持 flash_attention_2 实现高效推理

模型能力

图像理解

文本生成

视觉数学问题解答

多模态推理

使用案例

教育

视觉数学题解答

帮助学生解答包含图表和图像的数学问题

准确理解题目并给出解答

研究

多模态推理研究

用于视觉语言推理相关研究

🚀 OpenVLThinker-7B

OpenVLThinker-7B 是一个视觉语言推理模型，旨在处理多模态任务，尤其针对视觉数学问题解决进行了调优。

🚀 快速开始

模型信息

属性	详情
基础模型	Qwen/Qwen2.5-VL-7B-Instruct
许可证	apache-2.0
库名称	transformers
任务类型	图像文本到文本

💻 使用示例

基础用法

from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
import torch
from qwen_vl_utils import process_vision_info
import requests
from PIL import Image

# 1. Define model and processor names
model_name = "ydeng9/OpenVLThinker-7B"
processor_name = "Qwen/Qwen2.5-VL-7B-Instruct"

# 2. Load the OpenVLThinker-7B model and processor
device = "cuda:0" if torch.cuda.is_available() else "cpu"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
    device_map=device
)
processor = AutoProcessor.from_pretrained(processor_name)

# 3. Define a sample image URL and an instruction
image_url = "https://example.com/sample_image.jpg"  # replace with your image URL
instruction = "Example question"

# 4. Create a multimodal prompt using a chat message structure
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image_url},
            {"type": "text", "text": instruction},
        ],
    }
]

# 5. Generate a text prompt from the chat messages
text_prompt = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

# 6. Process image (and video) inputs from the messages
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text_prompt],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(device)

# 7. Generate the model's response (with specified generation parameters)
generated_ids = model.generate(
    **inputs,
    do_sample=True,
    max_new_tokens=2048,
    top_p=0.001,
    top_k=1,
    temperature=0.01,
    repetition_penalty=1.0,
)

# 8. Decode the generated tokens into human-readable text
generated_text = processor.batch_decode(
    generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False
)[0]

# 9. Print the generated response
print("Generated Response:")
print(generated_text)

引用信息

@misc{deng2025openvlthinker,
      title={OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement}, 
      author={Yihe Deng and Hritik Bansal and Fan Yin and Nanyun Peng and Wei Wang and Kai-Wei Chang},
      year={2025},
      eprint={2503.17352},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2503.17352}, 
}