SegVol开源医学体数据图像分割模型 - 支持多方式提示进行体积分割

首页

Segvol

由 yuxindu 开发

SegVol是一款通用且交互式的医学体数据图像分割模型，支持通过点提示、框提示和文本提示进行体积分割。

图像分割

Transformers

英语开源协议:MIT #交互式医学分割 #多模态提示 #3D体数据

下载量 16

发布时间 : 4/10/2024

模型简介

SegVol是一个用于医学体数据图像分割的基础模型，能够支持超过200种解剖结构的识别分割。通过在9万例未标注的CT扫描数据和6千例标注CT数据上进行训练，该模型具有强大的分割能力。

模型特点

多模态提示

支持通过点提示、框提示和文本提示进行交互式分割。

大规模训练数据

在9万例未标注的CT扫描数据和6千例标注CT数据上进行训练。

广泛解剖结构支持

能够识别和分割超过200种解剖结构。

3D医学图像处理

专门针对医学体数据（如CT扫描）进行优化。

模型能力

医学图像分割

3D体积分割

交互式分割

多模态提示分割

使用案例

医学影像分析

器官分割

自动分割CT扫描中的肝脏、肾脏、脾脏、胰腺等器官。

高精度的分割结果，可用于临床诊断和手术规划。

解剖结构识别

识别和分割超过200种不同的解剖结构。

提高医学影像分析的效率和准确性。

🚀 SegVol - 用于体积医学图像分割的通用交互式模型

SegVol 是一个用于体积医学图像分割的通用交互式模型。它接受点、框和文本提示，并输出体积分割结果。通过在 90k 个未标记的计算机断层扫描（CT）体积和 6k 个标记的 CT 上进行训练，这个基础模型支持对 200 多个解剖类别进行分割。

论文和代码已发布。

关键词：3D 医学 SAM，体积图像分割

image/jpeg

🚀 快速开始

🔧 环境要求

conda create -n segvol_transformers python=3.8
conda activate segvol_transformers

需要 pytorch v1.11.0（或更高版本）。使用以下命令安装关键依赖：

pip install 'monai[all]==0.9.0'
pip install einops==0.6.1
pip install transformers==4.18.0
pip install matplotlib

💻 测试脚本

from transformers import AutoModel, AutoTokenizer
import torch

# get device
device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu")

# load model
# IF you cannot connect to huggingface.co, you can download the repo and set from_pretrained path as the loacl dir path, replacing "yuxindu/segvol"
clip_tokenizer = AutoTokenizer.from_pretrained("yuxindu/segvol")
model = AutoModel.from_pretrained("yuxindu/segvol", trust_remote_code=True, test_mode=True)
model.model.text_encoder.tokenizer = clip_tokenizer
model.eval()
model.to(device)
print('model load done')

# set case path
# you can download this case from huggingface yuxindu/segvol files and versions
ct_path = 'path/to/Case_image_00001_0000.nii.gz'
gt_path = 'path/to/Case_label_00001.nii.gz'

# set categories, corresponding to the unique values(1, 2, 3, 4, ...) in ground truth mask
categories = ["liver", "kidney", "spleen", "pancreas"]

# generate npy data format
ct_npy, gt_npy = model.processor.preprocess_ct_gt(ct_path, gt_path, category=categories)
# IF you have download our 25 processed datasets, you can skip to here with the processed ct_npy, gt_npy files

# go through zoom_transform to generate zoomout & zoomin views
data_item = model.processor.zoom_transform(ct_npy, gt_npy)

# add batch dim manually
data_item['image'], data_item['label'], data_item['zoom_out_image'], data_item['zoom_out_label'] = \
data_item['image'].unsqueeze(0).to(device), data_item['label'].unsqueeze(0).to(device), data_item['zoom_out_image'].unsqueeze(0).to(device), data_item['zoom_out_label'].unsqueeze(0).to(device)

# take liver as the example
cls_idx = 0

# text prompt
text_prompt = [categories[cls_idx]]

# point prompt
point_prompt, point_prompt_map = model.processor.point_prompt_b(data_item['zoom_out_label'][0][cls_idx], device=device)   # inputs w/o batch dim, outputs w batch dim

# bbox prompt
bbox_prompt, bbox_prompt_map = model.processor.bbox_prompt_b(data_item['zoom_out_label'][0][cls_idx], device=device)   # inputs w/o batch dim, outputs w batch dim

print('prompt done')

# segvol test forward
# use_zoom: use zoom-out-zoom-in
# point_prompt_group: use point prompt
# bbox_prompt_group: use bbox prompt
# text_prompt: use text prompt
logits_mask = model.forward_test(image=data_item['image'],
      zoomed_image=data_item['zoom_out_image'],
      # point_prompt_group=[point_prompt, point_prompt_map],
      bbox_prompt_group=[bbox_prompt, bbox_prompt_map],
      text_prompt=text_prompt,
      use_zoom=False
      )

# cal dice score
dice = model.processor.dice_score(logits_mask[0][0], data_item['label'][0][cls_idx])
print(dice)

# save prediction as nii.gz file
save_path='./Case_preds_00001.nii.gz'
model.processor.save_preds(ct_path, save_path, logits_mask[0][0], 
                           start_coord=data_item['foreground_start_coord'], 
                           end_coord=data_item['foreground_end_coord'])
print('done')

💻 训练脚本

from transformers import AutoModel, AutoTokenizer
import torch

# get device
device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu")

# load model
# IF you cannot connect to huggingface.co, you can download the repo and set from_pretrained path as the loacl dir path, replacing "yuxindu/segvol"
clip_tokenizer = AutoTokenizer.from_pretrained("yuxindu/segvol")
model = AutoModel.from_pretrained("yuxindu/segvol", trust_remote_code=True, test_mode=False)
model.model.text_encoder.tokenizer = clip_tokenizer
model.train()
model.to(device)
print('model load done')

# set case path
# you can download this case from huggingface yuxindu/segvol files and versions
ct_path = 'path/to/Case_image_00001_0000.nii.gz'
gt_path = 'path/to/Case_label_00001.nii.gz'

# set categories, corresponding to the unique values(1, 2, 3, 4, ...) in ground truth mask
categories = ["liver", "kidney", "spleen", "pancreas"]

# generate npy data format
ct_npy, gt_npy = model.processor.preprocess_ct_gt(ct_path, gt_path, category=categories)
# IF you have download our 25 processed datasets, you can skip to here with the processed ct_npy, gt_npy files

# go through train transform
data_item = model.processor.train_transform(ct_npy, gt_npy)

# training example
# add batch dim manually
image, gt3D = data_item["image"].unsqueeze(0).to(device), data_item["label"].unsqueeze(0).to(device) # add batch dim

loss_step_avg = 0
for cls_idx in range(len(categories)):
    # optimizer.zero_grad()
    organs_cls = categories[cls_idx]
    labels_cls = gt3D[:, cls_idx]
    print(image.shape, organs_cls, labels_cls.shape)
    loss = model.forward_train(image, train_organs=organs_cls, train_labels=labels_cls)
    loss_step_avg += loss.item()
    loss.backward()
    # optimizer.step()

loss_step_avg /= len(categories)
print(f'AVG loss {loss_step_avg}')

# save ckpt
model.save_pretrained('./ckpt')