Qwen2.5-7B-Instruct开源大语言模型 - 免费部署助力文本生成与推理任务

首页

Chinese Text Correction 7b

由 shibing624 开发

Qwen2.5-7B-Instruct 是一个基于 Qwen2.5 架构的 7B 参数规模的中文指令微调大语言模型，适用于文本生成和推理任务。

大型语言模型

Transformers

中文开源协议:Apache-2.0 #中文文本纠错 #指令微调 #高精度语义理解

下载量 522

发布时间 : 10/12/2024

模型简介

该模型主要用于中文文本生成和推理任务，支持文本纠错等应用场景。

模型特点

中文指令微调

针对中文指令进行了优化，能够更好地理解和执行中文任务。

文本纠错能力

支持中文文本纠错任务，能够识别和修正文本中的错误。

大语言模型

基于 7B 参数规模的大语言模型，具备强大的文本生成和理解能力。

模型能力

文本生成

文本纠错

指令理解

使用案例

文本纠错

中文文本纠错

识别并修正中文文本中的语法、拼写和用词错误。

能够有效提升文本的准确性和可读性。

文本生成

中文文本生成

根据给定的提示生成连贯、流畅的中文文本。

生成的文本符合上下文逻辑，具有较高的可读性。

🚀 中文文本纠错模型

本项目提供的中文文本纠错模型，可用于拼写纠错和语法纠错，能有效提升文本的准确性和规范性。

🚀 快速开始

使用`pycorrector`调用模型

本项目开源在pycorrector项目：pycorrector，可支持大模型微调后用于文本纠错，通过如下命令调用：

安装依赖包：

pip install -U pycorrector

from pycorrector.gpt.gpt_corrector import GptCorrector

if __name__ == '__main__':
    error_sentences = [
        '真麻烦你了。希望你们好好的跳无',
        '少先队员因该为老人让坐',
        '机七学习是人工智能领遇最能体现智能的一个分知',
        '一只小鱼船浮在平净的河面上',
        '我的家乡是有明的渔米之乡',
    ]
    m = GptCorrector("shibing624/chinese-text-correction-7b")

    batch_res = m.correct_batch(error_sentences)
    for i in batch_res:
        print(i)
        print()

使用`HuggingFace Transformers`调用模型

若不使用 pycorrector，可以按如下方式使用模型：

首先，将输入数据传入transformer模型，然后得到生成的句子。

安装依赖包：

pip install transformers

# pip install transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
checkpoint = "shibing624/chinese-text-correction-7b"

device = "cuda" # for GPU usage or "cpu" for CPU usage
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint).to(device)

input_content = "文本纠错：\n少先队员因该为老人让坐。"

messages = [{"role": "user", "content": input_content}]
input_text=tokenizer.apply_chat_template(messages, tokenize=False)

print(input_text)

inputs = tokenizer.encode(input_text, return_tensors="pt").to(device)
outputs = model.generate(inputs, max_new_tokens=1024, temperature=0, do_sample=False, repetition_penalty=1.08)

print(tokenizer.decode(outputs[0]))

输出结果：

少先队员应该为老人让座。

✨ 主要特性

多类型纠错：支持拼写纠错、语法纠错，涵盖音似、形似、多字、少字等多种错误类型。
多方式调用：既可以通过pycorrector项目调用，也能使用HuggingFace Transformers直接调用。
多模型可选：提供不同规模的模型，如chinese-text-correction-1.5b、chinese-text-correction-7b等，满足不同场景需求。

📦 安装指南

使用`pycorrector`

pip install -U pycorrector

使用`HuggingFace Transformers`

pip install transformers

💻 使用示例

基础用法

# 使用pycorrector进行文本纠错
from pycorrector.gpt.gpt_corrector import GptCorrector

if __name__ == '__main__':
    error_sentences = [
        '真麻烦你了。希望你们好好的跳无',
        '少先队员因该为老人让坐',
        '机七学习是人工智能领遇最能体现智能的一个分知',
        '一只小鱼船浮在平净的河面上',
        '我的家乡是有明的渔米之乡',
    ]
    m = GptCorrector("shibing624/chinese-text-correction-7b")

    batch_res = m.correct_batch(error_sentences)
    for i in batch_res:
        print(i)
        print()

高级用法

# 使用HuggingFace Transformers直接调用模型进行文本纠错
# pip install transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
checkpoint = "shibing624/chinese-text-correction-7b"

device = "cuda" # for GPU usage or "cpu" for CPU usage
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint).to(device)

input_content = "文本纠错：\n少先队员因该为老人让坐。"

messages = [{"role": "user", "content": input_content}]
input_text=tokenizer.apply_chat_template(messages, tokenize=False)

print(input_text)

inputs = tokenizer.encode(input_text, return_tensors="pt").to(device)
outputs = model.generate(inputs, max_new_tokens=1024, temperature=0, do_sample=False, repetition_penalty=1.08)

print(tokenizer.decode(outputs[0]))

📚 详细文档

模型列表

模型名称	基础模型	下载链接
chinese-text-correction-1.5b	Qwen/Qwen2.5-1.5B-Instruct	🤗 Hugging Face
chinese-text-correction-1.5b-lora	Qwen/Qwen2.5-1.5B-Instruct	🤗 Hugging Face
chinese-text-correction-7b	Qwen/Qwen2.5-7B-Instruct	🤗 Hugging Face
chinese-text-correction-7b-lora	Qwen/Qwen2.5-7B-Instruct	🤗 Hugging Face

评估结果

评估指标：F1
CSC(Chinese Spelling Correction)：拼写纠错模型，表示模型可以处理音似、形似、语法等长度对齐的错误纠正。
CTC(CHinese Text Correction)：文本纠错模型，表示模型支持拼写、语法等长度对齐的错误纠正，还可以处理多字、少字等长度不对齐的错误纠正。
GPU：Tesla V100，显存 32 GB

模型名称	模型链接	基础模型	平均得分	SIGHAN - 2015得分	EC - LAW得分	MCSC得分	GPU/CPU	QPS
Kenlm - CSC	shibing624/chinese-kenlm-klm	kenlm	0.3409	0.3147	0.3763	0.3317	CPU	9
Mengzi - T5 - CSC	shibing624/mengzi-t5-base-chinese-correction	mengzi - t5 - base	0.3984	0.7758	0.3156	0.1039	GPU	214
ERNIE - CSC	PaddleNLP/ernie-csc	PaddlePaddle/ernie - 1.0 - base - zh	0.4353	0.8383	0.3357	0.1318	GPU	114
MacBERT - CSC	shibing624/macbert4csc-base-chinese	hfl/chinese - macbert - base	0.3993	0.8314	0.1610	0.2055	GPU	224
ChatGLM3 - 6B - CSC	shibing624/chatglm3-6b-csc-chinese-lora	THUDM/chatglm3 - 6b	0.4538	0.6572	0.4369	0.2672	GPU	3
Qwen2.5 - 1.5B - CTC	shibing624/chinese-text-correction-1.5b	Qwen/Qwen2.5 - 1.5B - Instruct	0.6802	0.3032	0.7846	0.9529	GPU	6
Qwen2.5 - 7B - CTC	shibing624/chinese-text-correction-7b	Qwen/Qwen2.5 - 7B - Instruct	0.8225	0.4917	0.9798	0.9959	GPU	3

模型文件组成

shibing624/chinese-text-correction-7b
|-- added_tokens.json
|-- config.json
|-- generation_config.json
|-- merges.txt
|-- model.safetensors
|-- model.safetensors.index.json
|-- README.md
|-- special_tokens_map.json
|-- tokenizer_config.json
|-- tokenizer.json
`-- vocab.json

训练参数

训练轮数（num_epochs）：8
批次大小（batch_size）：2
训练步数（steps）：36000
评估损失（eval_loss）：0.12
基础模型（base model）：Qwen/Qwen2.5 - 7B - Instruct
训练数据（train data）：shibing624/chinese_text_correction
训练时间（train time）：10 天
评估损失曲线：
训练损失曲线：

训练数据集

中文纠错数据集

数据：shibing624/chinese_text_correction

训练参考

如果需要训练Qwen的纠错模型，请参考https://github.com/shibing624/pycorrector 或者 https://github.com/shibing624/MedicalGPT

🔧 技术细节

本模型基于Qwen系列基础模型进行微调，使用特定的训练数据和训练参数，以提升在中文文本纠错任务上的性能。通过F1指标进行评估，在不同的测试数据集上表现良好。

📄 许可证

本项目采用apache - 2.0许可证。

📖 引用

@software{pycorrector,
  author = {Xu Ming},
  title = {pycorrector: Implementation of language model finetune},
  year = {2024},
  url = {https://github.com/shibing624/pycorrector},
}