PVT-Tiny-224 Open-Source Image Classification Model - Free Deployment for Accurately Completing Image Classification Tasks

Pvt Tiny 224

Developed by Xrenya

Pyramid Vision Transformer (PVT) is a vision model based on transformer architecture, specifically designed for image classification tasks.

Image Classification

Transformers

Open Source License:Apache-2.0 #Image Classification #Pyramid Structure #Convolution-Free Backbone

Downloads 25

Release Time : 3/25/2023

Model Overview

This model is pretrained and fine-tuned on the ImageNet-1K dataset, capable of classifying images into 1000 categories. It adopts a pyramid structure to reduce computational costs, making it suitable for dense prediction tasks.

Model Features

Pyramid Structure

Uses a progressive shrinking pyramid to reduce computational costs and improve efficiency in processing large feature maps.

Transformer Encoder

Based on transformer architecture, captures global image information through self-attention mechanisms.

CLS Token Classification

Uses the [CLS] token as a holistic representation of the image, facilitating classification tasks.

Model Capabilities

Image Classification

Feature Extraction

Use Cases

Computer Vision

Image Classification

Classifies input images into 1000 ImageNet categories.

Performs well on the ImageNet-1K dataset.

Feature Extraction

Extracts image features for downstream tasks.

🚀 Pyramid Vision Transformer (tiny-sized model)

The Pyramid Vision Transformer (PVT) is a pre - trained model on ImageNet - 1K for image classification tasks, offering efficient feature extraction and classification capabilities.

🚀 Quick Start

The Pyramid Vision Transformer (PVT) model is pre - trained on ImageNet - 1K (1 million images, 1000 classes) at a resolution of 224x224 and fine - tuned on ImageNet 2012 (1 million images, 1,000 classes) at the same resolution. It was introduced in the paper Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions by Wenhai Wang, Enze Xie, Xiang Li, Deng - Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, Ling Shao and first released in this repository.

Disclaimer: The team releasing PVT did not write a model card for this model, so this model card has been written by Rinat S. [@Xrenya].

✨ Features

Progressive Shrinking Pyramid: Unlike ViT models, PVT uses a progressive shrinking pyramid to reduce computations of large feature maps at each stage.
Inner Representation Learning: Through pre - training, the model learns an inner representation of images, which can be used for downstream tasks.

💻 Usage Examples

Basic Usage

Here is how to use this model to classify an image of the COCO 2017 dataset into one of the 1,000 ImageNet classes:

from transformers import PvtImageProcessor, PvtForImageClassification
from PIL import Image
import requests

url = 'http://images.cocodataset.org/val2017/000000039769.jpg'
image = Image.open(requests.get(url, stream=True).raw)

processor = PvtImageProcessor.from_pretrained('Zetatech/pvt-tiny-224')
model = PvtForImageClassification.from_pretrained('Zetatech/pvt-tiny-224')

inputs = processor(images=image, return_tensors="pt")
outputs = model(**inputs)
logits = outputs.logits
# model predicts one of the 1000 ImageNet classes
predicted_class_idx = logits.argmax(-1).item()
print("Predicted class:", model.config.id2label[predicted_class_idx])

For more code examples, we refer to the documentation.

📚 Documentation

Model description

The Pyramid Vision Transformer (PVT) is a transformer encoder model (BERT - like) pretrained on ImageNet - 1k (also referred to as ILSVRC2012), a dataset comprising 1 million images and 1,000 classes, also at resolution 224x224.

Images are presented to the model as a sequence of variable - size patches, which are linearly embedded. One also adds a [CLS] token to the beginning of a sequence to use it for classification tasks. One also adds absolute position embeddings before feeding the sequence to the layers of the Transformer encoder.

By pre - training the model, it learns an inner representation of images that can then be used to extract features useful for downstream tasks: if you have a dataset of labeled images for instance, you can train a standard classifier by placing a linear layer on top of the pre - trained encoder. One typically places a linear layer on top of the [CLS] token, as the last hidden state of this token can be seen as a representation of an entire image.

Intended uses & limitations

You can use the raw model for image classification. See the model hub to look for fine - tuned versions on a task that interests you.

🔧 Technical Details

Training data

The ViT model was pretrained on [ImageNet - 1k](http://www.image - net.org/challenges/LSVRC/2012/), a dataset consisting of 1 million images and 1k classes.

Training procedure

Preprocessing

The exact details of preprocessing of images during training/validation can be found here.

Images are resized/rescaled to the same resolution (224x224) and normalized across the RGB channels with mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225).

BibTeX entry and citation info

@inproceedings{wang2021pyramid,
  title={Pyramid vision transformer: A versatile backbone for dense prediction without convolutions},
  author={Wang, Wenhai and Xie, Enze and Li, Xiang and Fan, Deng - Ping and Song, Kaitao and Liang, Ding and Lu, Tong and Luo, Ping and Shao, Ling},
  booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
  pages={568--578},
  year={2021}
}

📄 License

This model is licensed under the Apache - 2.0 license.

Property	Details
Model Type	Pyramid Vision Transformer (tiny - sized model)
Training Data	ImageNet - 1k

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご