Pix2Struct-textcaps-base Open-Source Vision-Language Model - Free for Image Captioning Generation Tasks

Pix2struct Textcaps Base

Developed by google

Pix2Struct is a vision-language understanding model that processes image-to-text tasks through pre-training and fine-tuning, particularly suitable for image caption generation.

Image-to-Text

Transformers

Supports Multiple LanguagesOpen Source License:Apache-2.0 #Image Caption Generation #Multilingual Visual Understanding #HTML Structure Parsing

Downloads 3,888

Release Time : 3/1/2023

Model Overview

Pix2Struct is an image encoder-text decoder model trained on image-text pairs, suitable for various tasks such as image caption generation and visual question answering.

Model Features

Multi-domain Adaptability

Excels in multiple tasks across four major domains: documents, illustrations, user interfaces, and natural images.

Variable Resolution Input

Supports variable resolution input representations, adapting to images of different sizes.

Flexible Language-Vision Integration

Language prompts such as questions can be directly rendered on input images, enabling more flexible input integration.

Model Capabilities

Image Caption Generation

Visual Question Answering

OCR Recognition

Language Modeling

Use Cases

Image Understanding

Image Caption Generation

Generate natural language descriptions for input images.

Produces accurate and fluent image captions.

Visual Question Answering

Answer natural language questions about image content.

Provides accurate answers related to the image content.

Document Processing

Document Image to Text

Convert document images into structured text.

Extracts text content from documents while preserving structure.

🚀 Pix2Struct - Finetuned on TextCaps

Pix2Struct is an image encoder - text decoder model trained on image - text pairs, applicable to tasks like image captioning and visual question answering.

🚀 Quick Start

Pix2Struct is a powerful image - to - text model. This section will guide you through using the model, including conversion and running methods.

✨ Features

Versatile Applications: Can handle various tasks related to visually - situated language, such as image captioning and visual question answering.
Innovative Pretraining: Pretrained by learning to parse masked screenshots of web pages into simplified HTML, which provides rich and diverse pretraining data.
Flexible Input Representation: Introduces a variable - resolution input representation and a more flexible integration of language and vision inputs.

📦 Installation

Converting from T5x to huggingface

You can use the convert_pix2struct_checkpoint_to_pytorch.py script as follows:

python convert_pix2struct_checkpoint_to_pytorch.py --t5x_checkpoint_path PATH_TO_T5X_CHECKPOINTS --pytorch_dump_path PATH_TO_SAVE

If you are converting a large model, run:

python convert_pix2struct_checkpoint_to_pytorch.py --t5x_checkpoint_path PATH_TO_T5X_CHECKPOINTS --pytorch_dump_path PATH_TO_SAVE --use-large

Once saved, you can push your converted model with the following snippet:

from transformers import Pix2StructForConditionalGeneration, Pix2StructProcessor

model = Pix2StructForConditionalGeneration.from_pretrained(PATH_TO_SAVE)
processor = Pix2StructProcessor.from_pretrained(PATH_TO_SAVE)

model.push_to_hub("USERNAME/MODEL_NAME")
processor.push_to_hub("USERNAME/MODEL_NAME")

💻 Usage Examples

Basic Usage

Running the model in full precision on CPU

import requests
from PIL import Image
from transformers import Pix2StructForConditionalGeneration, Pix2StructProcessor

url = "https://www.ilankelman.org/stopsigns/australia.jpg"
image = Image.open(requests.get(url, stream=True).raw)

model = Pix2StructForConditionalGeneration.from_pretrained("google/pix2struct-textcaps-base")
processor = Pix2StructProcessor.from_pretrained("google/pix2struct-textcaps-base")

# image only
inputs = processor(images=image, return_tensors="pt")

predictions = model.generate(**inputs)
print(processor.decode(predictions[0], skip_special_tokens=True))
>>> A stop sign is on a street corner.

Advanced Usage

Running the model in full precision on GPU

import requests
from PIL import Image
from transformers import Pix2StructForConditionalGeneration, Pix2StructProcessor

url = "https://www.ilankelman.org/stopsigns/australia.jpg"
image = Image.open(requests.get(url, stream=True).raw)

model = Pix2StructForConditionalGeneration.from_pretrained("google/pix2struct-textcaps-base").to("cuda")
processor = Pix2StructProcessor.from_pretrained("google/pix2struct-textcaps-base")

# image only
inputs = processor(images=image, return_tensors="pt").to("cuda")

predictions = model.generate(**inputs)
print(processor.decode(predictions[0], skip_special_tokens=True))
>>> A stop sign is on a street corner.

Running the model in half precision on GPU

import requests
import torch

from PIL import Image
from transformers import Pix2StructForConditionalGeneration, Pix2StructProcessor

url = "https://www.ilankelman.org/stopsigns/australia.jpg"
image = Image.open(requests.get(url, stream=True).raw)

model = Pix2StructForConditionalGeneration.from_pretrained("google/pix2struct-textcaps-base", torch_dtype=torch.bfloat16).to("cuda")
processor = Pix2StructProcessor.from_pretrained("google/pix2struct-textcaps-base")

# image only
inputs = processor(images=image, return_tensors="pt").to("cuda", torch.bfloat16)

predictions = model.generate(**inputs)
print(processor.decode(predictions[0], skip_special_tokens=True))
>>> A stop sign is on a street corner.

Use different sequence length

This model has been trained on a sequence length of 2048. You can try to reduce the sequence length for a more memory - efficient inference but you may observe some performance degradation for small sequence length (<512). Just pass max_patches when calling the processor:

inputs = processor(images=image, return_tensors="pt", max_patches=512)

Conditional generation

You can also pre - pend some input text to perform conditional generation:

import requests
from PIL import Image
from transformers import Pix2StructForConditionalGeneration, Pix2StructProcessor

url = "https://www.ilankelman.org/stopsigns/australia.jpg"
image = Image.open(requests.get(url, stream=True).raw)
text = "A picture of"

model = Pix2StructForConditionalGeneration.from_pretrained("google/pix2struct-textcaps-base")
processor = Pix2StructProcessor.from_pretrained("google/pix2struct-textcaps-base")

# image only
inputs = processor(images=image, text=text, return_tensors="pt")

predictions = model.generate(**inputs)
print(processor.decode(predictions[0], skip_special_tokens=True))
>>> A picture of a stop sign that says yes.

📚 Documentation

TL;DR

Pix2Struct is an image encoder - text decoder model that is trained on image - text pairs for various tasks, including image captionning and visual question answering. The full list of available models can be found on the Table 1 of the paper:

Table 1 - paper

The abstract of the model states that:

Visually - situated language is ubiquitous—sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domainspecific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image - to - text model for purely visual language understanding, which can be finetuned on tasks containing visually - situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, image captioning. In addition to the novel pretraining strategy, we introduce a variable - resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state - of - the - art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.

📄 License

This model is licensed under apache - 2.0.

🔧 Technical Details

The model is pretrained by learning to parse masked screenshots of web pages into simplified HTML. This unique pretraining strategy allows the model to handle a wide variety of visually - situated language tasks. It also introduces a variable - resolution input representation and a more flexible integration of language and vision inputs, enabling better performance in different scenarios.

🤝 Contribution

This model was originally contributed by Kenton Lee, Mandar Joshi et al. and added to the Hugging Face ecosystem by Younes Belkada.

📖 Citation

If you want to cite this work, please consider citing the original paper:

@misc{https://doi.org/10.48550/arxiv.2210.03347,
  doi = {10.48550/ARXIV.2210.03347},
  
  url = {https://arxiv.org/abs/2210.03347},
  
  author = {Lee, Kenton and Joshi, Mandar and Turc, Iulia and Hu, Hexiang and Liu, Fangyu and Eisenschlos, Julian and Khandelwal, Urvashi and Shaw, Peter and Chang, Ming - Wei and Toutanova, Kristina},
  
  keywords = {Computation and Language (cs.CL), Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  
  title = {Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding},
  
  publisher = {arXiv},
  
  year = {2022},
  
  copyright = {Creative Commons Attribution 4.0 International}
}

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご