wav2vec2-large-xlsr-53-swedish开源模型 - 支持16kHz，精准识别瑞典语语音

首页

Wav2vec2 Large Xlsr 53 Swedish

由 KBLab 开发

基于facebook/wav2vec2-large-xlsr-53框架微调的瑞典语自动语音识别模型，支持16kHz采样率的语音输入

语音识别其他开源协议:Apache-2.0 #瑞典语语音识别 #低词错误率(WER14.3%)#XLSR-53微调

下载量 30.51k

发布时间 : 3/2/2022

模型简介

这是一个专门针对瑞典语优化的自动语音识别(ASR)模型，基于大规模XLSR-53架构，在瑞典NST听写语料库和通用语音库上进行了微调。

模型特点

高性能瑞典语识别

在通用语音库瑞典语测试集上达到14.3%的词错误率和4.93%的字符错误率

多阶段训练

经过预训练、增量训练和最终微调三个阶段优化

无需语言模型

可直接使用，无需额外语言模型支持

模型能力

瑞典语语音识别

音频转文本

语音处理

使用案例

语音转写

广播内容转录

将瑞典语广播节目自动转写为文本

语音指令识别

识别瑞典语语音命令

语音辅助技术

无障碍应用

为听障人士提供实时字幕服务

🚀 Wav2Vec2-Large-XLSR-53-瑞典语模型

本项目是对 facebook/wav2vec2-large-xlsr-53 模型进行瑞典语微调后的成果，使用了 NST Swedish Dictation 数据集。使用该模型时，请确保语音输入的采样率为 16kHz。

注意：为获得最佳性能，建议使用我们更新的模型 wav2vec2-large-voxrex-swedish。

🚀 快速开始

此模型可直接使用（无需语言模型），以下是使用示例：

import torch
import torchaudio
from datasets import load_dataset
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor

test_dataset = load_dataset("common_voice", "sv-SE", split="test[:2%]")

processor = Wav2Vec2Processor.from_pretrained("KBLab/wav2vec2-large-xlsr-53-swedish")
model = Wav2Vec2ForCTC.from_pretrained("KBLab/wav2vec2-large-xlsr-53-swedish")

resampler = torchaudio.transforms.Resample(48_000, 16_000)

# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
    speech_array, sampling_rate = torchaudio.load(batch["path"])
    batch["speech"] = resampler(speech_array).squeeze().numpy()

    return batch

test_dataset = test_dataset.map(speech_file_to_array_fn)
inputs = processor(test_dataset["speech"][:2], sampling_rate=16_000, return_tensors="pt", padding=True)

with torch.no_grad():
    logits = model(inputs.input_values, attention_mask=inputs.attention_mask).logits

predicted_ids = torch.argmax(logits, dim=-1)

print("Prediction:", processor.batch_decode(predicted_ids))
print("Reference:", test_dataset["sentence"][:2])

✨ 主要特性

基于 facebook/wav2vec2-large-xlsr-53 模型进行瑞典语微调。
可直接用于瑞典语的自动语音识别任务。

💻 使用示例

基础用法

import torch
import torchaudio
from datasets import load_dataset
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor

test_dataset = load_dataset("common_voice", "sv-SE", split="test[:2%]")

processor = Wav2Vec2Processor.from_pretrained("KBLab/wav2vec2-large-xlsr-53-swedish")
model = Wav2Vec2ForCTC.from_pretrained("KBLab/wav2vec2-large-xlsr-53-swedish")

resampler = torchaudio.transforms.Resample(48_000, 16_000)

# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
    speech_array, sampling_rate = torchaudio.load(batch["path"])
    batch["speech"] = resampler(speech_array).squeeze().numpy()

    return batch

test_dataset = test_dataset.map(speech_file_to_array_fn)
inputs = processor(test_dataset["speech"][:2], sampling_rate=16_000, return_tensors="pt", padding=True)

with torch.no_grad():
    logits = model(inputs.input_values, attention_mask=inputs.attention_mask).logits

predicted_ids = torch.argmax(logits, dim=-1)

print("Prediction:", processor.batch_decode(predicted_ids))
print("Reference:", test_dataset["sentence"][:2])

📚 详细文档

评估

该模型可在 Common Voice 的瑞典语测试数据上进行评估，示例代码如下：

import torch
import torchaudio
from datasets import load_dataset, load_metric
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import re

test_dataset = load_dataset("common_voice", "sv-SE", split="test")
wer = load_metric("wer")

processor = Wav2Vec2Processor.from_pretrained("KBLab/wav2vec2-large-xlsr-53-swedish")
model = Wav2Vec2ForCTC.from_pretrained("KBLab/wav2vec2-large-xlsr-53-swedish")
model.to("cuda")

chars_to_ignore_regex = '[,?.!\\-;:"“]'
resampler = torchaudio.transforms.Resample(48_000, 16_000)

# Preprocessing the datasets.
# We need to read the aduio files as arrays
def speech_file_to_array_fn(batch):
    batch["sentence"] = re.sub(chars_to_ignore_regex, '', batch["sentence"]).lower()
    speech_array, sampling_rate = torchaudio.load(batch["path"])
    batch["speech"] = resampler(speech_array).squeeze().numpy()

    return batch

test_dataset = test_dataset.map(speech_file_to_array_fn)

# Preprocessing the datasets.
# We need to read the aduio files as arrays
def evaluate(batch):
    inputs = processor(batch["speech"], sampling_rate=16_000, return_tensors="pt", padding=True)

    with torch.no_grad():
        logits = model(inputs.input_values.to("cuda"), attention_mask=inputs.attention_mask.to("cuda")).logits

    pred_ids = torch.argmax(logits, dim=-1)
    batch["pred_strings"] = processor.batch_decode(pred_ids)

    return batch

result = test_dataset.map(evaluate, batched=True, batch_size=8)

print("WER: {:2f}".format(100 * wer.compute(predictions=result["pred_strings"], references=result["sentence"])))
print("CER: {:2f}".format(100 * wer.compute(predictions=[" ".join(list(entry)) for entry in result["pred_strings"]], references=[" ".join(list(entry)) for entry in result["sentence"]])))

字错率（WER）：14.298610% 字符错误率（CER）：4.925294%

训练

模型的训练过程如下：

首先，使用包含来自各个广播电台的 1000 小时瑞典语语音语料，对 XLSR 模型进行 50 个周期的预训练。
其次，使用 NST Swedish Dictation 和 Common Voice 数据集进行微调。
最后，仅使用 Common Voice 数据集进行最终微调。训练使用了 Fairseq 脚本。

📄 许可证

本项目采用 Apache 2.0 许可证。

📦 数据集和指标

属性	详情
数据集	common_voice、KTH/nst
评估指标	字错率（wer）、字符错误率（cer）

📦 模型信息

模型名称	任务	数据集	评估指标
XLSR Wav2Vec2 Swedish by KBLab	语音识别（automatic-speech-recognition）	Common Voice sv-SE	测试字错率（Test WER）：14.298610 测试字符错误率（Test CER）：4.925294