免费部署Gemma 3-27B IT量化模型，支持视觉文本输入输出，推理高效

首页

Gemma 3 27b It Quantized.w4a16

由 RedHatAI 开发

这是google/gemma-3-27b-it的量化版本，支持视觉-文本输入和文本输出，通过权重量化和激活量化优化，可使用vLLM进行高效推理。

图像生成文本

Transformers

#多模态理解 #高效推理 #INT4量化

下载量 302

发布时间 : 6/4/2025

模型简介

基于google/gemma-3-27b-it的量化模型，支持多模态输入和文本生成，适用于需要高效推理的场景。

模型特点

高效量化

采用INT4权重量化和FP16激活量化，显著减少模型大小和推理资源需求。

多模态支持

支持视觉-文本输入，能够处理图像和文本结合的复杂任务。

高效推理

优化后可通过vLLM进行高效部署和推理，提升处理速度。

模型能力

多模态理解

文本生成

图像内容分析

使用案例

内容理解与生成

图像内容描述

分析图像内容并生成描述性文本

准确识别图像中的主要元素和场景

多模态对话

结合图像和文本输入进行智能对话

提供与图像内容相关的连贯回答

🚀 gemma-3-27b-it-quantized.w4a16

这是 google/gemma-3-27b-it 的量化版本，该模型支持视觉 - 文本输入，输出为文本。通过权重量化和激活量化等优化，可使用 vLLM 进行高效推理。

🚀 快速开始

本模型可使用 vLLM 后端进行高效部署，示例代码如下：

from vllm import LLM, SamplingParams
from vllm.assets.image import ImageAsset
from transformers import AutoProcessor

# Define model name once
model_name = "RedHatAI/gemma-3-27b-it-quantized.w4a16"

# Load image and processor
image = ImageAsset("cherry_blossom").pil_image.convert("RGB")
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)

# Build multimodal prompt
chat = [
    {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "What is the content of this image?"}]},
    {"role": "assistant", "content": []}
]
prompt = processor.apply_chat_template(chat, add_generation_prompt=True)

# Initialize model
llm = LLM(model=model_name, trust_remote_code=True)

# Run inference
inputs = {"prompt": prompt, "multi_modal_data": {"image": [image]}}
outputs = llm.generate(inputs, SamplingParams(temperature=0.2, max_tokens=64))

# Display result
print("RESPONSE:", outputs[0].outputs[0].text)

vLLM 还支持与 OpenAI 兼容的服务，更多详情请参阅文档。

✨ 主要特性

模型概述

模型架构：google/gemma-3-27b-it
- 输入：视觉 - 文本
- 输出：文本
模型优化：
- 权重量化：INT4
- 激活量化：FP16
发布日期：2025 年 4 月 6 日
版本：1.0
模型开发者：RedHatAI

模型优化

本模型通过将 google/gemma-3-27b-it 的权重量化为 INT4 数据类型获得，可使用 vLLM >= 0.8.0 进行推理。

💻 使用示例

基础用法

from vllm import LLM, SamplingParams
from vllm.assets.image import ImageAsset
from transformers import AutoProcessor

# Define model name once
model_name = "RedHatAI/gemma-3-27b-it-quantized.w4a16"

# Load image and processor
image = ImageAsset("cherry_blossom").pil_image.convert("RGB")
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)

# Build multimodal prompt
chat = [
    {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "What is the content of this image?"}]},
    {"role": "assistant", "content": []}
]
prompt = processor.apply_chat_template(chat, add_generation_prompt=True)

# Initialize model
llm = LLM(model=model_name, trust_remote_code=True)

# Run inference
inputs = {"prompt": prompt, "multi_modal_data": {"image": [image]}}
outputs = llm.generate(inputs, SamplingParams(temperature=0.2, max_tokens=64))

# Display result
print("RESPONSE:", outputs[0].outputs[0].text)

🔧 技术细节

模型创建

本模型使用 llm-compressor 创建，代码如下：

模型创建代码

import base64
from io import BytesIO
import torch
from datasets import load_dataset
from transformers import AutoProcessor, Gemma3ForConditionalGeneration
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.transformers import oneshot


# Load model.
model_id = "google/gemma-3-27b-it"
model = Gemma3ForConditionalGeneration.from_pretrained(
    model_id,
    device_map="auto",
    torch_dtype="auto",
)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)

# Oneshot arguments
DATASET_ID = "neuralmagic/calibration"
DATASET_SPLIT = {"LLM": "train[:1024]"}
NUM_CALIBRATION_SAMPLES = 1024
MAX_SEQUENCE_LENGTH = 2048

# Load dataset and preprocess.
ds = load_dataset(DATASET_ID, split=DATASET_SPLIT)
ds = ds.shuffle(seed=42)

dampening_frac=0.07

def data_collator(batch):
    assert len(batch) == 1, "Only batch size of 1 is supported for calibration"
    item = batch[0]
    collated = {}
    import torch


    for key, value in item.items():
        if isinstance(value, torch.Tensor):
            collated[key] = value.unsqueeze(0)
        elif isinstance(value, list) and isinstance(value[0][0], int):
            # Handle tokenized inputs like input_ids, attention_mask
            collated[key] = torch.tensor(value)
        elif isinstance(value, list) and isinstance(value[0][0], float):
            # Handle possible float sequences
            collated[key] = torch.tensor(value)
        elif isinstance(value, list) and isinstance(value[0][0], torch.Tensor):
            # Handle batched image data (e.g., pixel_values as [C, H, W])
            collated[key] = torch.stack(value)  # -> [1, C, H, W]
        elif isinstance(value, torch.Tensor):
            collated[key] = value
        else:
            print(f"[WARN] Unrecognized type in collator for key={key}, type={type(value)}")
    
    return collated
   


# Recipe
recipe = [
    GPTQModifier(
        targets="Linear",
        ignore=["re:.*lm_head.*", "re:.*embed_tokens.*", "re:vision_tower.*", "re:multi_modal_projector.*"],
        sequential_update=True,
        sequential_targets=["Gemma3DecoderLayer"],
        dampening_frac=dampening_frac,
    )
]

SAVE_DIR=f"{model_id.split('/')[1]}-quantized.w4a16"

# Perform oneshot
oneshot(
    model=model,
    tokenizer=model_id,
    dataset=ds,
    recipe=recipe,
    max_seq_length=MAX_SEQUENCE_LENGTH,
    num_calibration_samples=NUM_CALIBRATION_SAMPLES,
    trust_remote_code_model=True,
    data_collator=data_collator,
    output_dir=SAVE_DIR
)

模型评估

本模型使用 lm_evaluation_harness 进行 OpenLLM v1 文本基准测试，评估命令如下：

评估命令

OpenLLM v1

lm_eval \
  --model vllm \
  --model_args pretrained="<model_name>",dtype=auto,add_bos_token=True,max_model_len=4096,tensor_parallel_size=<n>,gpu_memory_utilization=0.8,enable_chunked_prefill=True,trust_remote_code=True,enforce_eager=True \
  --tasks openllm \
  --batch_size auto

准确率

类别	指标	google/gemma-3-27b-it	RedHatAI/gemma-3-27b-it-quantized.w8a8	恢复率 (%)
OpenLLM V1	ARC Challenge	72.53%	72.35%	99.76%
OpenLLM V1	GSM8K	92.12%	91.66%	99.51%
OpenLLM V1	Hellaswag	85.78%	84.97%	99.06%
OpenLLM V1	MMLU	77.53%	76.77%	99.02%
OpenLLM V1	Truthfulqa (mc2)	62.20%	62.57%	100.59%
OpenLLM V1	Winogrande	79.40%	79.79%%	100.50%
OpenLLM V1	平均得分	78.26%	78.02%	99.70%
视觉评估	MMMU (val)	50.89%	51.78%	101.75%
视觉评估	ChartQA	72.16%	72.20%	100.06%
视觉评估	平均得分	61.53%	61.99%	100.90%