Pixtral-12B開源多模態模型 - 免費部署實現圖像理解與描述任務

首頁

Pixtral 12b

由mgoin開發

Pixtral-12B 是一個與 transformers 庫兼容的多模態模型，能夠處理圖像和文本輸入並生成文本輸出，適用於圖像理解和描述任務。

圖像生成文本

Transformers

#多圖像理解 #圖文對話生成 #大語言模型集成

下載量 1,943

發布時間 : 10/18/2024

模型概述

Pixtral-12B 是一個基於 Mistral 架構的多模態模型，支持圖像和文本的聯合處理，能夠生成高質量的圖像描述和回答相關問題。

模型特點

多模態處理

能夠同時處理圖像和文本輸入，生成連貫的文本輸出。

高質量圖像描述

能夠生成詳細且準確的圖像描述，包括場景、物體和情感分析。

聊天模板支持

支持使用聊天模板格式化聊天曆史記錄，便於多輪對話。

模型能力

圖像描述

多模態問答

場景分析

物體識別

使用案例

圖像理解

圖像描述生成

輸入一張或多張圖像，模型生成詳細的描述文本。

生成包含場景、物體和情感分析的詳細描述。

多模態問答

結合圖像和文本提問，模型生成相關回答。

能夠根據圖像內容回答相關問題，提供上下文相關的信息。

自然語言處理

聊天機器人

支持多輪對話，結合圖像和文本進行交互。

生成連貫且上下文相關的回答。

🚀 圖像文本轉文本模型 `pixtral`

pixtral 是與 transformers 庫兼容的模型檢查點。它能夠處理圖像和文本輸入，並生成相應的文本輸出，為圖像理解和描述任務提供了強大的支持。

🚀 快速開始

在使用 pixtral 模型之前，請確保從源代碼安裝 transformers 庫，或者等待 v4.45 版本發佈。

基礎用法

from PIL import Image
from transformers import AutoProcessor, LlavaForConditionalGeneration
model_id = "mistral-community/pixtral-12b"
model = LlavaForConditionalGeneration.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)

IMG_URLS = [
"https://picsum.photos/id/237/400/300", 
"https://picsum.photos/id/231/200/300", 
"https://picsum.photos/id/27/500/500",
"https://picsum.photos/id/17/150/600",
]
PROMPT = "<s>[INST]Describe the images.\n[IMG][IMG][IMG][IMG][/INST]"

inputs = processor(text=PROMPT, images=IMG_URLS, return_tensors="pt").to("cuda")
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

運行上述代碼後，你應該會得到類似以下的輸出：

"""
Describe the images.
Sure, let's break down each image description:

1. **Image 1:**
   - **Description:** A black dog with a glossy coat is sitting on a wooden floor. The dog has a focused expression and is looking directly at the camera.
   - **Details:** The wooden floor has a rustic appearance with visible wood grain patterns. The dog's eyes are a striking color, possibly brown or amber, which contrasts with its black fur.

2. **Image 2:**
   - **Description:** A scenic view of a mountainous landscape with a winding road cutting through it. The road is surrounded by lush green vegetation and leads to a distant valley.
   - **Details:** The mountains are rugged with steep slopes, and the sky is clear, indicating good weather. The winding road adds a sense of depth and perspective to the image.

3. **Image 3:**
   - **Description:** A beach scene with waves crashing against the shore. There are several people in the water and on the beach, enjoying the waves and the sunset.
   - **Details:** The waves are powerful, creating a dynamic and lively atmosphere. The sky is painted with hues of orange and pink from the setting sun, adding a warm glow to the scene.

4. **Image 4:**
   - **Description:** A garden path leading to a large tree with a bench underneath it. The path is bordered by well-maintained grass and flowers.
   - **Details:** The path is made of small stones or gravel, and the tree provides a shaded area with the bench invitingly placed beneath it. The surrounding area is lush and green, suggesting a well-kept garden.

Each image captures a different scene, from a close-up of a dog to expansive natural landscapes, showcasing various elements of nature and human interaction with it.
"""

高級用法

你還可以使用聊天模板來格式化 Pixtral 的聊天曆史記錄。確保 processor 的 images 參數包含的圖像順序與聊天中出現的順序一致，以便模型理解每個圖像的位置。

from PIL import Image
from transformers import AutoProcessor, LlavaForConditionalGeneration
model_id = "mistral-community/pixtral-12b"
model = LlavaForConditionalGeneration.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)

url_dog = "https://picsum.photos/id/237/200/300"
url_mountain = "https://picsum.photos/seed/picsum/200/300"

chat = [
    {
      "role": "user", "content": [
        {"type": "text", "content": "Can this animal"}, 
        {"type": "image"}, 
        {"type": "text", "content": "live here?"}, 
        {"type": "image"}
      ]
    }
]

prompt = processor.apply_chat_template(chat)
inputs = processor(text=prompt, images=[url_dog, url_mountain], return_tensors="pt").to(model.device)
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

運行上述代碼後，你應該會得到類似以下的輸出：

Can this animallive here?Certainly! Here are some details about the images you provided:

### First Image
- **Description**: The image shows a black dog lying on a wooden surface. The dog has a curious expression with its head tilted slightly to one side.
- **Details**: The dog appears to be a young puppy with soft, shiny fur. Its eyes are wide and alert, and it has a playful demeanor.
- **Context**: This image could be used to illustrate a pet-friendly environment or to showcase the dog's personality.

### Second Image
- **Description**: The image depicts a serene landscape with a snow-covered hill in the foreground. The sky is painted with soft hues of pink, orange, and purple, indicating a sunrise or sunset.
- **Details**: The hill is covered in a blanket of pristine white snow, and the horizon meets the sky in a gentle curve. The scene is calm and peaceful.
- **Context**: This image could be used to represent tranquility, natural beauty, or a winter wonderland.

### Combined Context
If you're asking whether the dog can "live here," referring to the snowy landscape, it would depend on the breed and its tolerance to cold weather. Some breeds, like Huskies or Saint Bernards, are well-adapted to cold environments, while others might struggle. The dog in the first image appears to be a breed that might prefer warmer climates.

Would you like more information on any specific aspect?