SpatialBot-3B开源视觉语言模型 - 精准解析深度图，高效执行高级任务

首页

Spatialbot 3B

由 RussRobin 开发

SpatialBot是一款具备空间理解与推理能力的视觉语言模型，能精准解析深度图并执行高级任务。

文本生成图像

Transformers

英语#深度图解析 #空间推理 #多模态VLM

下载量 301

发布时间 : 7/17/2024

模型简介

基于Phi-2和SigLIP架构开发的融合版视觉语言模型，在常规视觉语言任务及空间理解基准测试中表现优异。

模型特点

空间理解能力

能够精准解析深度图并进行空间推理

多模态处理

同时处理视觉和语言输入，实现跨模态理解

高效架构

基于Phi-2和SigLIP的高效架构设计

模型能力

深度图解析

空间推理

视觉问答

多模态理解

使用案例

空间理解

深度值查询

从深度图中读取指定坐标点的深度值

精确返回深度数值

空间关系推理

分析场景中物体的空间位置关系

生成准确的空间描述

🚀 SpatialBot - 具备空间理解能力的视觉语言模型

SpatialBot是一款具备空间理解和推理能力的视觉语言模型（VLM），它能够精准理解深度图，并利用这些信息执行高级任务。在这个Hugging Face仓库中，我们提供了基于Phi - 2和SigLIP的SpatialBot - 3B合并模型。该模型在一般的VLM任务以及像SpatialBench这样的空间理解基准测试中表现出色。

🚀 快速开始

⚠️ 重要提示

我们在2024年8月28日更新了仓库和快速启动代码。如果您在此日期之前下载了模型和代码，请进行更新。

📦 安装指南

首先，安装依赖项：

pip install torch transformers accelerate pillow numpy

💻 使用示例

基础用法

运行模型的代码如下：

import torch
import transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
from PIL import Image
import warnings
import numpy as np

# disable some warnings
transformers.logging.set_verbosity_error()
transformers.logging.disable_progress_bar()
warnings.filterwarnings('ignore')

# set device
device = 'cuda'  # or cpu

model_name = 'RussRobin/SpatialBot-3B'
offset_bos = 0

# create model
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16, # float32 for cpu
    device_map='auto',
    trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(
    model_name,
    trust_remote_code=True)

# text prompt
prompt = 'What is the depth value of point <0.5,0.2>? Answer directly from depth map.'
text = f"A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions. USER: <image 1>\n<image 2>\n{prompt} ASSISTANT:"
text_chunks = [tokenizer(chunk).input_ids for chunk in text.split('<image 1>\n<image 2>\n')]
input_ids = torch.tensor(text_chunks[0] + [-201] + [-202] + text_chunks[1][offset_bos:], dtype=torch.long).unsqueeze(0).to(device)

image1 = Image.open('rgb.jpg')
image2 = Image.open('depth.png')

channels = len(image2.getbands())
if channels == 1:
    img = np.array(image2)
    height, width = img.shape
    three_channel_array = np.zeros((height, width, 3), dtype=np.uint8)
    three_channel_array[:, :, 0] = (img // 1024) * 4
    three_channel_array[:, :, 1] = (img // 32) * 8
    three_channel_array[:, :, 2] = (img % 32) * 8
    image2 = Image.fromarray(three_channel_array, 'RGB')

image_tensor = model.process_images([image1,image2], model.config).to(dtype=model.dtype, device=device)

# generate
output_ids = model.generate(
    input_ids,
    images=image_tensor,
    max_new_tokens=100,
    use_cache=True,
    repetition_penalty=1.0 # increase this to avoid chattering
)[0]

print(tokenizer.decode(output_ids[input_ids.shape[1]:], skip_special_tokens=True).strip())