Pixtral-12bオープンソースマルチモーダルモデル - 画像テキストを無料で処理し、詳細な説明を生成

Home

Pixtral 12b

Developed by mistral-community

PixtralはMistralアーキテクチャを基にしたマルチモーダルモデルで、画像とテキスト入力を処理し、詳細なテキスト記述を生成できます。

画像生成テキスト

Transformers

Open Source License:Apache-2.0 #マルチイメージテキスト生成 #視覚-言語インタラクション #複雑なシーン理解

Downloads 31.93k

Release Time : 9/13/2024

Model Overview

Pixtralは12Bパラメータのマルチモーダルモデルで、画像からテキストへのタスク向けに設計されており、画像内容を理解し詳細な記述や質問への回答を生成できます。

Model Features

マルチモーダル能力

画像とテキスト入力を同時に処理し、一貫性のあるテキスト出力を生成できます。

大規模パラメータ

12Bパラメータの規模により、強力な理解と生成能力を備えています。

柔軟な入力形式

URLまたはローカルパスを通じて画像をロードでき、チャットテンプレートで入力をフォーマットできます。

Model Capabilities

画像記述生成

マルチイメージ分析

画像質問応答

マルチモーダル対話

Use Cases

コンテンツ生成

画像記述生成

単一または複数の画像に対して詳細なテキスト記述を生成します。

画像の詳細、背景、感情的なニュアンスを含む記述テキストを生成します。

質問応答システム

画像関連質問回答

画像内容に基づいてユーザーの質問に回答します。

画像内容に関連する正確な回答と説明を提供します。

🚀 トランスフォーマー

Transformers互換のpixtralチェックポイントです。ソースからインストールするか、v4.45のリリースを待つことをおすすめします！

📄 ライセンス

このライブラリはApache-2.0ライセンスの下で提供されています。

💻 使用例

基本的な使用法

from PIL import Image
from transformers import AutoProcessor, LlavaForConditionalGeneration
model_id = "mistral-community/pixtral-12b"
model = LlavaForConditionalGeneration.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)

IMG_URLS = [
"https://picsum.photos/id/237/400/300", 
"https://picsum.photos/id/231/200/300", 
"https://picsum.photos/id/27/500/500",
"https://picsum.photos/id/17/150/600",
]
PROMPT = "<s>[INST]Describe the images.\n[IMG][IMG][IMG][IMG][/INST]"

inputs = processor(text=PROMPT, images=IMG_URLS, return_tensors="pt").to("cuda")
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

以下のような出力が得られるはずです：


"""
Describe the images.
Sure, let's break down each image description:

1. **Image 1:**
   - **Description:** A black dog with a glossy coat is sitting on a wooden floor. The dog has a focused expression and is looking directly at the camera.
   - **Details:** The wooden floor has a rustic appearance with visible wood grain patterns. The dog's eyes are a striking color, possibly brown or amber, which contrasts with its black fur.

2. **Image 2:**
   - **Description:** A scenic view of a mountainous landscape with a winding road cutting through it. The road is surrounded by lush green vegetation and leads to a distant valley.
   - **Details:** The mountains are rugged with steep slopes, and the sky is clear, indicating good weather. The winding road adds a sense of depth and perspective to the image.

3. **Image 3:**
   - **Description:** A beach scene with waves crashing against the shore. There are several people in the water and on the beach, enjoying the waves and the sunset.
   - **Details:** The waves are powerful, creating a dynamic and lively atmosphere. The sky is painted with hues of orange and pink from the setting sun, adding a warm glow to the scene.

4. **Image 4:**
   - **Description:** A garden path leading to a large tree with a bench underneath it. The path is bordered by well-maintained grass and flowers.
   - **Details:** The path is made of small stones or gravel, and the tree provides a shaded area with the bench invitingly placed beneath it. The surrounding area is lush and green, suggesting a well-kept garden.

Each image captures a different scene, from a close-up of a dog to expansive natural landscapes, showcasing various elements of nature and human interaction with it.
"""

高度な使用法

チャットテンプレートを使用して、Pixtralのチャット履歴を整形することもできます。processorのimages引数に、チャット内で表示される順序で画像を含めるようにしてください。これにより、モデルが各画像の位置を理解できます。

以下は、同じメッセージ内でテキストと複数の画像を交互に配置した例です：

from PIL import Image
from transformers import AutoProcessor, LlavaForConditionalGeneration
model_id = "mistral-community/pixtral-12b"
model = LlavaForConditionalGeneration.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)

url_dog = "https://picsum.photos/id/237/200/300"
url_mountain = "https://picsum.photos/seed/picsum/200/300"

chat = [
    {
      "role": "user", "content": [
        {"type": "text", "content": "Can this animal"}, 
        {"type": "image"}, 
        {"type": "text", "content": "live here?"}, 
        {"type": "image"}
      ]
    }
]

prompt = processor.apply_chat_template(chat)
inputs = processor(text=prompt, images=[url_dog, url_mountain], return_tensors="pt").to(model.device)
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

Transformers v4.48以降では、会話履歴に画像のURLまたはローカルパスを渡すことができ、チャットテンプレートが残りの処理を行います。チャットテンプレートは画像を自動的に読み込み、torch.Tensor形式の入力を返します。これを直接model.generate()に渡すことができます。

chat = [
    {
      "role": "user", "content": [
        {"type": "text", "content": "Can this animal"}, 
        {"type": "image", "url": url_dog}, 
        {"type": "text", "content": "live here?"}, 
        {"type": "image", "url" : url_mountain}
      ]
    }
]

inputs = processor.apply_chat_template(chat, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt").to(model.device)
generate_ids = model.generate(**inputs, max_new_tokens=500)
output = processor.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]

以下のような出力が得られるはずです：

Can this animallive here?Certainly! Here are some details about the images you provided:

### First Image
- **Description**: The image shows a black dog lying on a wooden surface. The dog has a curious expression with its head tilted slightly to one side.
- **Details**: The dog appears to be a young puppy with soft, shiny fur. Its eyes are wide and alert, and it has a playful demeanor.
- **Context**: This image could be used to illustrate a pet-friendly environment or to showcase the dog's personality.

### Second Image
- **Description**: The image depicts a serene landscape with a snow-covered hill in the foreground. The sky is painted with soft hues of pink, orange, and purple, indicating a sunrise or sunset.
- **Details**: The hill is covered in a blanket of pristine white snow, and the horizon meets the sky in a gentle curve. The scene is calm and peaceful.
- **Context**: This image could be used to represent tranquility, natural beauty, or a winter wonderland.

### Combined Context
If you're asking whether the dog can "live here," referring to the snowy landscape, it would depend on the breed and its tolerance to cold weather. Some breeds, like Huskies or Saint Bernards, are well-adapted to cold environments, while others might struggle. The dog in the first image appears to be a breed that might prefer warmer climates.

Would you like more information on any specific aspect?

入力の間隔が崩れているように見えるかもしれませんが、これは表示のために特殊トークンをスキップしているためです。実際には、「Can this animal」と「live here」は画像トークンによって正しく区切られています。特殊トークンを含めてデコードすると、モデルが見ている内容を正確に確認できます！