Pixtral-12Bオープンソースマルチモーダルモデル - 無料でデプロイして画像理解と記述タスクを実現

ホーム

Pixtral 12b

mgoinによって開発

Pixtral-12Bは、transformersライブラリと互換性のあるマルチモーダルモデルで、画像とテキストの入力を処理し、テキスト出力を生成することができ、画像理解と記述タスクに適しています。

画像生成テキスト

Transformers

#多画像理解 #画像とテキストの対話生成 #大規模言語モデルの統合

ダウンロード数 1,943

リリース時間 : 10/18/2024

モデル概要

Pixtral-12Bは、Mistralアーキテクチャに基づくマルチモーダルモデルで、画像とテキストの統合処理をサポートし、高品質の画像記述を生成し、関連する質問に回答することができます。

モデル特徴

マルチモーダル処理

画像とテキストの入力を同時に処理し、首尾一貫したテキスト出力を生成することができます。

高品質の画像記述

シーン、物体、感情分析を含む詳細で正確な画像記述を生成することができます。

チャットテンプレートのサポート

チャットテンプレートを使用してチャット履歴を整形することをサポートし、複数回の対話を容易にします。

モデル能力

画像記述

マルチモーダル質問応答

シーン分析

物体認識

使用事例

画像理解

画像記述生成

1枚または複数枚の画像を入力すると、モデルは詳細な記述テキストを生成します。

シーン、物体、感情分析を含む詳細な記述を生成します。

マルチモーダル質問応答

画像とテキストを組み合わせた質問を行うと、モデルは関連する回答を生成します。

画像内容に基づいて関連する質問に回答し、文脈に関連する情報を提供することができます。

自然言語処理

チャットボット

複数回の対話をサポートし、画像とテキストを組み合わせて対話を行います。

首尾一貫した文脈に関連する回答を生成します。

🚀 画像テキストをテキストに変換するモデル `pixtral`

pixtral は transformers ライブラリと互換性のあるモデルチェックポイントです。画像とテキストの入力を処理し、対応するテキスト出力を生成することができ、画像理解と記述タスクに強力なサポートを提供します。

🚀 クイックスタート

pixtral モデルを使用する前に、ソースコードから transformers ライブラリをインストールするか、v4.45 バージョンのリリースを待つことを確認してください。

💻 使用例

基本的な使用法

from PIL import Image
from transformers import AutoProcessor, LlavaForConditionalGeneration
model_id = "mistral-community/pixtral-12b"
model = LlavaForConditionalGeneration.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)

IMG_URLS = [
"https://picsum.photos/id/237/400/300", 
"https://picsum.photos/id/231/200/300", 
"https://picsum.photos/id/27/500/500",
"https://picsum.photos/id/17/150/600",
]
PROMPT = "<s>[INST]Describe the images.\n[IMG][IMG][IMG][IMG][/INST]"

inputs = processor(text=PROMPT, images=IMG_URLS, return_tensors="pt").to("cuda")
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

上記のコードを実行すると、以下のような出力が得られるはずです。

"""
Describe the images.
Sure, let's break down each image description:

1. **Image 1:**
   - **Description:** A black dog with a glossy coat is sitting on a wooden floor. The dog has a focused expression and is looking directly at the camera.
   - **Details:** The wooden floor has a rustic appearance with visible wood grain patterns. The dog's eyes are a striking color, possibly brown or amber, which contrasts with its black fur.

2. **Image 2:**
   - **Description:** A scenic view of a mountainous landscape with a winding road cutting through it. The road is surrounded by lush green vegetation and leads to a distant valley.
   - **Details:** The mountains are rugged with steep slopes, and the sky is clear, indicating good weather. The winding road adds a sense of depth and perspective to the image.

3. **Image 3:**
   - **Description:** A beach scene with waves crashing against the shore. There are several people in the water and on the beach, enjoying the waves and the sunset.
   - **Details:** The waves are powerful, creating a dynamic and lively atmosphere. The sky is painted with hues of orange and pink from the setting sun, adding a warm glow to the scene.

4. **Image 4:**
   - **Description:** A garden path leading to a large tree with a bench underneath it. The path is bordered by well-maintained grass and flowers.
   - **Details:** The path is made of small stones or gravel, and the tree provides a shaded area with the bench invitingly placed beneath it. The surrounding area is lush and green, suggesting a well-kept garden.

Each image captures a different scene, from a close-up of a dog to expansive natural landscapes, showcasing various elements of nature and human interaction with it.
"""

高度な使用法

Pixtral のチャット履歴をフォーマットするために、チャットテンプレートを使用することもできます。processor の images パラメータに含まれる画像の順序が、チャットで表示される順序と一致するようにして、モデルが各画像の位置を理解できるようにしてください。

from PIL import Image
from transformers import AutoProcessor, LlavaForConditionalGeneration
model_id = "mistral-community/pixtral-12b"
model = LlavaForConditionalGeneration.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)

url_dog = "https://picsum.photos/id/237/200/300"
url_mountain = "https://picsum.photos/seed/picsum/200/300"

chat = [
    {
      "role": "user", "content": [
        {"type": "text", "content": "Can this animal"}, 
        {"type": "image"}, 
        {"type": "text", "content": "live here?"}, 
        {"type": "image"}
      ]
    }
]

prompt = processor.apply_chat_template(chat)
inputs = processor(text=prompt, images=[url_dog, url_mountain], return_tensors="pt").to(model.device)
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

上記のコードを実行すると、以下のような出力が得られるはずです。

Can this animallive here?Certainly! Here are some details about the images you provided:

### First Image
- **Description**: The image shows a black dog lying on a wooden surface. The dog has a curious expression with its head tilted slightly to one side.
- **Details**: The dog appears to be a young puppy with soft, shiny fur. Its eyes are wide and alert, and it has a playful demeanor.
- **Context**: This image could be used to illustrate a pet-friendly environment or to showcase the dog's personality.

### Second Image
- **Description**: The image depicts a serene landscape with a snow-covered hill in the foreground. The sky is painted with soft hues of pink, orange, and purple, indicating a sunrise or sunset.
- **Details**: The hill is covered in a blanket of pristine white snow, and the horizon meets the sky in a gentle curve. The scene is calm and peaceful.
- **Context**: This image could be used to represent tranquility, natural beauty, or a winter wonderland.

### Combined Context
If you're asking whether the dog can "live here," referring to the snowy landscape, it would depend on the breed and its tolerance to cold weather. Some breeds, like Huskies or Saint Bernards, are well-adapted to cold environments, while others might struggle. The dog in the first image appears to be a breed that might prefer warmer climates.

Would you like more information on any specific aspect?

⚠️ 重要提示

入力の間隔が乱れているように見えるかもしれませんが、これは表示のために特殊なマークをスキップしたためです。実際には、「Can this animal」と「live here」は画像マークで正しく区切られています。特殊マークを含めてデコードして、モデルが実際に見ている内容を確認することができます！