trocr-small-stage1 Open Source OCR Model - Suitable for Optical Character Recognition of Single-line Text Images

Trocr Small Stage1

Developed by microsoft

TrOCR is a Transformer-based pre-trained optical character recognition model that adopts an encoder-decoder architecture, suitable for OCR tasks on single-line text images.

Image-to-Text

Transformers

#Single-line Text OCR #Transformer Architecture #Image to Text

Downloads 3,713

Release Time : 3/2/2022

Model Overview

The TrOCR model combines an image Transformer encoder with a text Transformer decoder, capable of converting text in images into readable text content.

Model Features

Transformer-based Architecture

Utilizes advanced Transformer architecture for processing images and text, combining the strengths of DeiT and UniLM.

Pre-trained Model

Provides pre-trained weights that can be directly used for OCR tasks or as a base model for fine-tuning.

Single-line Text Recognition

Specifically optimized for optical character recognition tasks on single-line text images.

Model Capabilities

Image to Text

Optical Character Recognition

Single-line Text Recognition

Use Cases

Document Digitization

Scanned Document Recognition

Convert scanned document images into editable text content

High-precision text conversion results

Automated Processing

Form Processing

Automatically recognize and extract text information from forms

Improves data processing efficiency

🚀 TrOCR (small-sized model, pre-trained only)

TrOCR is a pre-trained only model for optical character recognition (OCR). It offers a novel approach using Transformer architectures for OCR tasks.

🚀 Quick Start

TrOCR is a pre - trained model introduced in the paper TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models by Li et al. and first released in this repository. You can utilize it for OCR on single text - line images.

✨ Features

Encoder - Decoder Architecture: The TrOCR model is an encoder - decoder model. It uses an image Transformer as the encoder and a text Transformer as the decoder.
Initialization from Pre - trained Models: The image encoder is initialized from the weights of DeiT, and the text decoder is initialized from the weights of UniLM.
Patch - based Image Input: Images are presented to the model as a sequence of fixed - size patches (resolution 16x16), which are linearly embedded.

📚 Documentation

Model description

The TrOCR model is an encoder - decoder model, consisting of an image Transformer as encoder, and a text Transformer as decoder. The image encoder was initialized from the weights of DeiT, while the text decoder was initialized from the weights of UniLM.

Images are presented to the model as a sequence of fixed - size patches (resolution 16x16), which are linearly embedded. One also adds absolute position embeddings before feeding the sequence to the layers of the Transformer encoder. Next, the Transformer text decoder autoregressively generates tokens.

Intended uses & limitations

You can use the raw model for optical character recognition (OCR) on single text - line images. See the model hub to look for fine - tuned versions on a task that interests you.

💻 Usage Examples

Basic Usage

from transformers import TrOCRProcessor, VisionEncoderDecoderModel
from PIL import Image
import requests
import torch

# load image from the IAM database
url = 'https://fki.tic.heia - fr.ch/static/img/a01 - 122 - 02 - 00.jpg'
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")

processor = TrOCRProcessor.from_pretrained('microsoft/trocr - small - stage1')
model = VisionEncoderDecoderModel.from_pretrained('microsoft/trocr - small - stage1')

# training
pixel_values = processor(image, return_tensors="pt").pixel_values  # Batch size 1
decoder_input_ids = torch.tensor([[model.config.decoder.decoder_start_token_id]])
outputs = model(pixel_values=pixel_values, decoder_input_ids=decoder_input_ids)

BibTeX entry and citation info

@misc{li2021trocr,
      title={TrOCR: Transformer - based Optical Character Recognition with Pre - trained Models}, 
      author={Minghao Li and Tengchao Lv and Lei Cui and Yijuan Lu and Dinei Florencio and Cha Zhang and Zhoujun Li and Furu Wei},
      year={2021},
      eprint={2109.10282},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご