VideoMAE Open-source Video Classification Model - Free Deployment for Video Classification on Kinetics Dataset

Videomae Large Finetuned Kinetics

Developed by MCG-NJU

VideoMAE is a self-supervised video pre-training model based on masked autoencoder, fine-tuned on the Kinetics-400 dataset for video classification tasks.

Video Processing

Transformers

#Video Self-Supervised Learning #High-Precision Action Recognition #Masked Autoencoder

Downloads 4,657

Release Time : 8/2/2022

Model Overview

This model is pre-trained in a self-supervised manner and fine-tuned under supervision on Kinetics-400, capable of classifying videos into 400 possible categories.

Model Features

Self-Supervised Pre-Training

Uses masked autoencoder (MAE) method for self-supervised video pre-training, with high data efficiency

Strong Video Understanding Capability

Demonstrates excellent video classification performance after fine-tuning on the Kinetics-400 dataset

Transformer Architecture

Based on Vision Transformer architecture, effectively processes video sequence data

Model Capabilities

Video Classification

Video Feature Extraction

Video Content Understanding

Use Cases

Video Content Analysis

Video Classification

Classifies videos into one of the 400 Kinetics-400 categories

Achieves 84.7% top-1 accuracy on the Kinetics-400 test set

Video Content Understanding

Extracts high-level feature representations of videos

🚀 VideoMAE (large-sized model, fine-tuned on Kinetics-400)

VideoMAE is a model pre - trained self - supervised for 1600 epochs and fine - tuned on Kinetics - 400. It offers an effective way for video classification.

🚀 Quick Start

VideoMAE model was pre - trained for 1600 epochs in a self - supervised way and then fine - tuned in a supervised way on Kinetics - 400. It was introduced in the paper VideoMAE: Masked Autoencoders are Data - Efficient Learners for Self - Supervised Video Pre - Training by Tong et al. and first released in this repository.

Disclaimer: The team releasing VideoMAE did not write a model card for this model, so this model card has been written by the Hugging Face team.

✨ Features

Model description

VideoMAE is an extension of Masked Autoencoders (MAE) to video. The model's architecture is quite similar to that of a standard Vision Transformer (ViT), with a decoder on top for predicting pixel values for masked patches.

Videos are presented to the model as a sequence of fixed - size patches (resolution 16x16), which are linearly embedded. A [CLS] token is added to the beginning of a sequence for classification tasks. Fixed sinus/cosinus position embeddings are also added before feeding the sequence to the layers of the Transformer encoder.

Through pre - training, the model learns an inner representation of videos, which can be used to extract features useful for downstream tasks. For example, if you have a dataset of labeled videos, you can train a standard classifier by placing a linear layer on top of the pre - trained encoder. Usually, a linear layer is placed on top of the [CLS] token, as the last hidden state of this token can be seen as a representation of an entire video.

Intended uses & limitations

You can use the raw model for video classification into one of the 400 possible Kinetics - 400 labels.

💻 Usage Examples

Basic Usage

Here is how to use this model to classify a video:

from transformers import VideoMAEImageProcessor, VideoMAEForVideoClassification
import numpy as np
import torch

video = list(np.random.randn(16, 3, 224, 224))

processor = VideoMAEImageProcessor.from_pretrained("MCG-NJU/videomae-large-finetuned-kinetics")
model = VideoMAEForVideoClassification.from_pretrained("MCG-NJU/videomae-large-finetuned-kinetics")

inputs = processor(video, return_tensors="pt")

with torch.no_grad():
  outputs = model(**inputs)
  logits = outputs.logits

predicted_class_idx = logits.argmax(-1).item()
print("Predicted class:", model.config.id2label[predicted_class_idx])

For more code examples, we refer to the documentation.

🔧 Technical Details

Evaluation results

This model obtains a top - 1 accuracy of 84.7 and a top - 5 accuracy of 96.5 on the test set of Kinetics - 400.

BibTeX entry and citation info

misc{https://doi.org/10.48550/arxiv.2203.12602,
  doi = {10.48550/ARXIV.2203.12602},
  url = {https://arxiv.org/abs/2203.12602},
  author = {Tong, Zhan and Song, Yibing and Wang, Jue and Wang, Limin},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {VideoMAE: Masked Autoencoders are Data - Efficient Learners for Self - Supervised Video Pre - Training},
  publisher = {arXiv},
  year = {2022},
  copyright = {Creative Commons Attribution 4.0 International}
}

📄 License

This model is licensed under "cc - by - nc - 4.0".

Property	Details
License	cc - by - nc - 4.0
Tags	vision, video - classification

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご