moonshine-tiny Open-source Automatic Speech Recognition Model - Free English Transcription for Resource-constrained Devices

Moonshine Tiny

Developed by UsefulSensors

The Moonshine model is an automatic speech recognition (ASR) model developed by Useful Sensors, focusing on efficient English speech transcription on resource-constrained devices.

Speech Recognition

Transformers

EnglishOpen Source License:MIT #Lightweight speech recognition #Real-time transcription #Low-resource deployment

Downloads 7,848

Release Time : 10/30/2024

Model Overview

The Moonshine model is a sequence-to-sequence automatic speech recognition system capable of transcribing English speech audio into English text. This model is particularly suitable for deployment on platforms with limited memory and computational resources.

Model Features

Efficient and lightweight

Optimized for resource-constrained devices, with a small parameter size but excellent performance

Real-time transcription

Supports real-time speech transcription, suitable for instant application scenarios

English-optimized

Specifically trained and optimized for English speech recognition

Model Capabilities

English speech recognition

Real-time speech transcription

Use Cases

Assistive tools

Real-time caption generation

Provides real-time captions for video conferences or live streams

Enhances accessibility and user experience

Voice command recognition

Implements voice control functionality on embedded devices

Enables touch-free device interaction

Content processing

Audio content transcription

Converts audio content such as podcasts and lectures into text

Facilitates content search and archiving

🚀 Moonshine

Moonshine is an automatic speech recognition (ASR) model developed by Useful Sensors. It can transcribe English speech audio into text, offering high accuracy on standard datasets. With different sizes and capabilities, it's suitable for platforms with limited memory and computational resources.

🚀 Quick Start

Moonshine is supported in Hugging Face 🤗 Transformers. To run the model, first install the Transformers library. For this example, we'll also install 🤗 Datasets to load toy audio dataset from the Hugging Face Hub, and 🤗 Accelerate to reduce the model loading time:

pip install --upgrade pip
pip install --upgrade transformers datasets[audio]

from transformers import MoonshineForConditionalGeneration, AutoProcessor
from datasets import load_dataset, Audio
import torch

device = "cuda:0" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32

model = MoonshineForConditionalGeneration.from_pretrained('UsefulSensors/moonshine-tiny').to(device).to(torch_dtype)
processor = AutoProcessor.from_pretrained('UsefulSensors/moonshine-tiny')

dataset = load_dataset("hf-internal-testing/librispeech_asr_dummy", "clean", split="validation")
dataset = dataset.cast_column("audio", Audio(processor.feature_extractor.sampling_rate))
sample = dataset[0]["audio"]

inputs = processor(
    sample["array"], 
    return_tensors="pt",
    sampling_rate=processor.feature_extractor.sampling_rate
)
inputs = inputs.to(device, torch_dtype)

# to avoid hallucination loops, we limit the maximum length of the generated text based expected number of tokens per second
token_limit_factor = 6.5 / processor.feature_extractor.sampling_rate  # Maximum of 6.5 tokens per second
seq_lens = inputs.attention_mask.sum(dim=-1)
max_length = int((seq_lens * token_limit_factor).max().item())

generated_ids = model.generate(**inputs, max_length=max_length)
print(processor.decode(generated_ids[0], skip_special_tokens=True))

✨ Features

High Accuracy: Exhibits greater accuracy on standard datasets over existing ASR systems of similar sizes.
Multiple Sizes: Available in different sizes (tiny and base) to meet various resource requirements.
Sequence-to-Sequence Architecture: Capable of handling speech recognition and translation tasks.

📦 Installation

To use Moonshine, you need to install the necessary libraries. Run the following commands:

pip install --upgrade pip
pip install --upgrade transformers datasets[audio]

💻 Usage Examples

Basic Usage

The above code in the "Quick Start" section demonstrates the basic usage of Moonshine. It shows how to load the model, process audio input, and generate transcription.

📚 Documentation

Model Details

The Moonshine models are trained for the speech recognition task, capable of transcribing English speech audio into English text. Useful Sensors developed the models to support their business direction of developing real time speech transcription products based on low cost hardware. There are 2 models of different sizes and capabilities, summarized in the following table.

Property	Details
Model Type	Sequence-to-sequence ASR (automatic speech recognition) and speech translation model
Release Date	October 2024
Paper & Samples	Paper / Blog

Size	Parameters	English-only model	Multilingual model
tiny	27 M	✓
base	61 M	✓

Model Use

Evaluated Use

The primary intended users of these models are AI developers that want to deploy English speech recognition systems in platforms that are severely constrained in memory capacity and computational resources. We recognize that once models are released, it is impossible to restrict access to only “intended” uses or to draw reasonable guidelines around what is or is not safe use.

The models are primarily trained and evaluated on English ASR task. They may exhibit additional capabilities, particularly if fine-tuned on certain tasks like voice activity detection, speaker classification, or speaker diarization but have not been robustly evaluated in these areas. We strongly recommend that users perform robust evaluations of the models in a particular context and domain before deploying them.

In particular, we caution against using Moonshine models to transcribe recordings of individuals taken without their consent or purporting to use these models for any kind of subjective classification. We recommend against use in high-risk domains like decision-making contexts, where flaws in accuracy can lead to pronounced flaws in outcomes. The models are intended to transcribe English speech, use of the model for classification is not only not evaluated but also not appropriate, particularly to infer human attributes.

Training Data

The models are trained on 200,000 hours of audio and the corresponding transcripts collected from the internet, as well as datasets openly available and accessible on HuggingFace. The open datasets used are listed in the the accompanying paper.

Performance and Limitations

Our evaluations show that, the models exhibit greater accuracy on standard datasets over existing ASR systems of similar sizes.

However, like any machine learning model, the predictions may include texts that are not actually spoken in the audio input (i.e. hallucination). We hypothesize that this happens because, given their general knowledge of language, the models combine trying to predict the next word in audio with trying to transcribe the audio itself.

In addition, the sequence-to-sequence architecture of the model makes it prone to generating repetitive texts, which can be mitigated to some degree by beam search and temperature scheduling but not perfectly. It is likely that this behavior and hallucinations may be worse for short audio segments, or segments where parts of words are cut off at the beginning or the end of the segment.

Broader Implications

We anticipate that Moonshine models’ transcription capabilities may be used for improving accessibility tools, especially for real-time transcription. The real value of beneficial applications built on top of Moonshine models suggests that the disparate performance of these models may have real economic implications.

There are also potential dual-use concerns that come with releasing Moonshine. While we hope the technology will be used primarily for beneficial purposes, making ASR technology more accessible could enable more actors to build capable surveillance technologies or scale up existing surveillance efforts, as the speed and accuracy allow for affordable automatic transcription and translation of large volumes of audio communication. Moreover, these models may have some capabilities to recognize specific individuals out of the box, which in turn presents safety concerns related both to dual use and disparate performance. In practice, we expect that the cost of transcription is not the limiting factor of scaling up surveillance projects.

🔧 Technical Details

The models are trained for the speech recognition task, aiming to transcribe English speech audio into English text. The sequence-to-sequence architecture is used, which combines the prediction of the next word in the audio with the transcription of the audio itself. However, this architecture may lead to issues such as hallucination and repetitive text generation.

📄 License

This project is licensed under the MIT license.

📖 Citation

If you benefit from our work, please cite us:

@misc{jeffries2024moonshinespeechrecognitionlive,
      title={Moonshine: Speech Recognition for Live Transcription and Voice Commands}, 
      author={Nat Jeffries and Evan King and Manjunath Kudlur and Guy Nicholson and James Wang and Pete Warden},
      year={2024},
      eprint={2410.15608},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2410.15608}, 
}

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご