Norbert3-base开源挪威语语言模型 - 支持两种挪威语的文本处理

首页

Norbert3 Base

由 ltg 开发

NorBERT 3 是新一代挪威语语言模型，基于BERT架构，支持书面挪威语（Bokmål）和新挪威语（Nynorsk）。

大型语言模型

Transformers

其他开源协议:Apache-2.0 #挪威语BERT变体 #掩码语言建模 #多方言支持

下载量 345

发布时间 : 3/2/2023

模型简介

NorBERT 3 是一个基于BERT架构的挪威语语言模型，主要用于自然语言处理任务，如文本分类、命名实体识别等。

模型特点

多语言支持

支持书面挪威语（Bokmål）和新挪威语（Nynorsk）两种挪威语变体。

多种规模版本

提供从超小版到大型版的多种参数规模，适应不同计算资源需求。

自定义封装器

需要从`modeling_norbert.py`加载自定义封装器，支持多种自然语言处理任务。

模型能力

文本分类

命名实体识别

问答系统

文本生成

语言理解

使用案例

自然语言处理

文本分类

用于挪威语文本的分类任务，如情感分析、主题分类等。

命名实体识别

识别挪威语文本中的命名实体，如人名、地名、组织名等。

问答系统

构建挪威语问答系统，回答用户提出的问题。

🚀 NorBERT 3 base

NorBERT 3 base是新一代NorBERT语言模型的官方版本，该模型在论文NorBench — A Benchmark for Norwegian Language Models中有所描述。请阅读论文以了解该模型的更多细节。

🚀 快速开始

本模型目前需要来自 modeling_norbert.py 的自定义包装器，因此你应该使用 trust_remote_code=True 来加载模型。

💻 使用示例

基础用法

import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

tokenizer = AutoTokenizer.from_pretrained("ltg/norbert3-base")
model = AutoModelForMaskedLM.from_pretrained("ltg/norbert3-base", trust_remote_code=True)

mask_id = tokenizer.convert_tokens_to_ids("[MASK]")
input_text = tokenizer("Nå ønsker de seg en[MASK] bolig.", return_tensors="pt")
output_p = model(**input_text)
output_text = torch.where(input_text.input_ids == mask_id, output_p.logits.argmax(-1), input_text.input_ids)

# should output: '[CLS] Nå ønsker de seg en ny bolig.[SEP]'
print(tokenizer.decode(output_text[0].tolist()))

目前实现了以下类：AutoModel、AutoModelMaskedLM、AutoModelForSequenceClassification、AutoModelForTokenClassification、AutoModelForQuestionAnswering 和 AutoModeltForMultipleChoice。

📚 详细文档

其他尺寸模型

生成式NorT5兄弟模型

📄 许可证

本项目采用 apache-2.0 许可证。

🔖 引用我们

如果你使用了本模型，请按照以下格式引用：

@inproceedings{samuel-etal-2023-norbench,
    title = "{N}or{B}ench {--} A Benchmark for {N}orwegian Language Models",
    author = "Samuel, David  and
      Kutuzov, Andrey  and
      Touileb, Samia  and
      Velldal, Erik  and
      {\O}vrelid, Lilja  and
      R{\o}nningstad, Egil  and
      Sigdel, Elina  and
      Palatkina, Anna",
    booktitle = "Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)",
    month = may,
    year = "2023",
    address = "T{\'o}rshavn, Faroe Islands",
    publisher = "University of Tartu Library",
    url = "https://aclanthology.org/2023.nodalida-1.61",
    pages = "618--633",
    abstract = "We present NorBench: a streamlined suite of NLP tasks and probes for evaluating Norwegian language models (LMs) on standardized data splits and evaluation metrics. We also introduce a range of new Norwegian language models (both encoder and encoder-decoder based). Finally, we compare and analyze their performance, along with other existing LMs, across the different benchmark tests of NorBench.",
}