Qwen2.5-1.5B-Instruct开源中文指令模型 - 支持文本生成与推理任务

首页

Chinese Text Correction 1.5b

由 shibing624 开发

Qwen2.5-1.5B-Instruct 是一个基于 Qwen2.5 架构的 15 亿参数的中文指令微调模型，适用于文本生成和推理任务。

大型语言模型

Transformers

中文开源协议:Apache-2.0 #中文文本纠错 #指令微调模型 #1.5B参数规模

下载量 1,085

发布时间 : 10/12/2024

模型简介

该模型主要用于中文文本生成和推理任务，支持多种自然语言处理应用，如文本纠错、问答系统等。

模型特点

中文指令微调

模型经过指令微调，能够更好地理解和执行中文指令任务。

文本生成能力

支持高质量的中文文本生成，适用于多种应用场景。

文本纠错

能够对中文文本进行纠错，提升文本质量。

模型能力

文本生成

文本纠错

问答系统

指令执行

使用案例

文本纠错

中文文本纠错

对中文文本中的语法、拼写错误进行纠正。

提升文本的准确性和可读性。

问答系统

中文问答

回答用户提出的中文问题。

提供准确且相关的答案。

🚀 中文文本纠错模型

本项目的中文文本纠错模型可用于拼写纠错、语法纠错，为中文文本的准确性提供有力支持。

🚀 快速开始

本项目开源在pycorrector项目：pycorrector，可支持大模型微调后用于文本纠错，通过如下命令调用：

使用`pycorrector`库

安装依赖包：

pip install -U pycorrector

from pycorrector.gpt.gpt_corrector import GptCorrector

if __name__ == '__main__':
    error_sentences = [
        '真麻烦你了。希望你们好好的跳无',
        '少先队员因该为老人让坐',
        '机七学习是人工智能领遇最能体现智能的一个分知',
        '一只小鱼船浮在平净的河面上',
        '我的家乡是有明的渔米之乡',
    ]
    m = GptCorrector("shibing624/chinese-text-correction-1.5b")

    batch_res = m.correct_batch(error_sentences)
    for i in batch_res:
        print(i)
        print()

使用HuggingFace Transformers库

若不使用 pycorrector，可以按以下方式使用模型：

首先，将输入数据传入transformer模型，然后获取生成的句子。

安装依赖包：

pip install transformers

# pip install transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
checkpoint = "shibing624/chinese-text-correction-1.5b"

device = "cuda" # for GPU usage or "cpu" for CPU usage
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint).to(device)

input_content = "文本纠错：\n少先队员因该为老人让坐。"

messages = [{"role": "user", "content": input_content}]
input_text=tokenizer.apply_chat_template(messages, tokenize=False)

print(input_text)

inputs = tokenizer.encode(input_text, return_tensors="pt").to(device)
outputs = model.generate(inputs, max_new_tokens=1024, temperature=0, do_sample=False, repetition_penalty=1.08)

print(tokenizer.decode(outputs[0]))

输出结果：

少先队员应该为老人让座。

✨ 主要特性

支持拼写纠错、语法纠错，可处理音似、形似、语法等长度对齐的错误纠正，还能处理多字、少字等长度不对齐的错误纠正。
提供多种不同规模的模型供选择，以满足不同场景的需求。

📦 安装指南

使用`pycorrector`库

pip install -U pycorrector

使用HuggingFace Transformers库

pip install transformers

💻 使用示例

基础用法

from pycorrector.gpt.gpt_corrector import GptCorrector

if __name__ == '__main__':
    error_sentences = [
        '真麻烦你了。希望你们好好的跳无',
        '少先队员因该为老人让坐',
        '机七学习是人工智能领遇最能体现智能的一个分知',
        '一只小鱼船浮在平净的河面上',
        '我的家乡是有明的渔米之乡',
    ]
    m = GptCorrector("shibing624/chinese-text-correction-1.5b")

    batch_res = m.correct_batch(error_sentences)
    for i in batch_res:
        print(i)
        print()

高级用法

# pip install transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
checkpoint = "shibing624/chinese-text-correction-1.5b"

device = "cuda" # for GPU usage or "cpu" for CPU usage
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint).to(device)

input_content = "文本纠错：\n少先队员因该为老人让坐。"

messages = [{"role": "user", "content": input_content}]
input_text=tokenizer.apply_chat_template(messages, tokenize=False)

print(input_text)

inputs = tokenizer.encode(input_text, return_tensors="pt").to(device)
outputs = model.generate(inputs, max_new_tokens=1024, temperature=0, do_sample=False, repetition_penalty=1.08)

print(tokenizer.decode(outputs[0]))

📚 详细文档

模型列表

模型名称	基础模型	下载链接
chinese-text-correction-1.5b	Qwen/Qwen2.5-1.5B-Instruct	🤗 Hugging Face
chinese-text-correction-1.5b-lora	Qwen/Qwen2.5-1.5B-Instruct	🤗 Hugging Face
chinese-text-correction-7b	Qwen/Qwen2.5-7B-Instruct	🤗 Hugging Face
chinese-text-correction-7b-lora	Qwen/Qwen2.5-7B-Instruct	🤗 Hugging Face

评估结果

评估指标：F1
CSC(Chinese Spelling Correction)：拼写纠错模型，表示模型可以处理音似、形似、语法等长度对齐的错误纠正
CTC(CHinese Text Correction)：文本纠错模型，表示模型支持拼写、语法等长度对齐的错误纠正，还可以处理多字、少字等长度不对齐的错误纠正
GPU：Tesla V100，显存 32 GB

模型名称	模型链接	基础模型	平均得分	SIGHAN - 2015	EC - LAW	MCSC	GPU/CPU	QPS
Kenlm - CSC	shibing624/chinese-kenlm-klm	kenlm	0.3409	0.3147	0.3763	0.3317	CPU	9
Mengzi - T5 - CSC	shibing624/mengzi-t5-base-chinese-correction	mengzi - t5 - base	0.3984	0.7758	0.3156	0.1039	GPU	214
ERNIE - CSC	PaddleNLP/ernie-csc	PaddlePaddle/ernie - 1.0 - base - zh	0.4353	0.8383	0.3357	0.1318	GPU	114
MacBERT - CSC	shibing624/macbert4csc-base-chinese	hfl/chinese - macbert - base	0.3993	0.8314	0.1610	0.2055	GPU	224
ChatGLM3 - 6B - CSC	shibing624/chatglm3-6b-csc-chinese-lora	THUDM/chatglm3 - 6b	0.4538	0.6572	0.4369	0.2672	GPU	3
Qwen2.5 - 1.5B - CTC	shibing624/chinese-text-correction-1.5b	Qwen/Qwen2.5 - 1.5B - Instruct	0.6802	0.3032	0.7846	0.9529	GPU	6
Qwen2.5 - 7B - CTC	shibing624/chinese-text-correction-7b	Qwen/Qwen2.5 - 7B - Instruct	0.8225	0.4917	0.9798	0.9959	GPU	3

模型文件组成

shibing624/chinese-text-correction-1.5b
|-- added_tokens.json
|-- config.json
|-- generation_config.json
|-- merges.txt
|-- model.safetensors
|-- model.safetensors.index.json
|-- README.md
|-- special_tokens_map.json
|-- tokenizer_config.json
|-- tokenizer.json
`-- vocab.json

训练参数

属性	详情
训练轮数	8
批次大小	4
训练步数	36000
评估损失	0.14
基础模型	Qwen/Qwen2.5 - 1.5B - Instruct
训练数据	shibing624/chinese_text_correction
训练时长	9天8小时
评估损失曲线
训练损失曲线

训练数据集

中文纠错数据集

数据：shibing624/chinese_text_correction

训练参考

如果需要训练Qwen的纠错模型，请参考https://github.com/shibing624/pycorrector 或者 https://github.com/shibing624/MedicalGPT

🔧 技术细节

评估指标

采用F1作为评估指标，全面衡量模型在不同数据集上的纠错性能。

模型类型

分为CSC和CTC两种类型，分别针对不同类型的中文文本错误进行处理。

训练环境

使用Tesla V100 GPU（显存32GB）进行训练，确保模型能够高效收敛。

📄 许可证

本项目采用apache - 2.0许可证。

📄 引用信息

@software{pycorrector,
  author = {Xu Ming},
  title = {pycorrector: Implementation of language model finetune},
  year = {2024},
  url = {https://github.com/shibing624/pycorrector},
}