Hubert_emotion开源语音情感识别模型 - 从音频中精准识别说话者情感状态

首页

Hubert Emotion

由 Rajaram1996 开发

基于Hubert架构的语音情感识别模型，能够从音频中识别说话者的情感状态。

音频分类

Transformers

#语音情感识别 #性别区分情感 #高精度分类

下载量 76

发布时间 : 3/2/2022

模型简介

该模型使用Hubert架构进行训练，专门用于语音情感分类任务。它可以识别多种情感状态，如悲伤、恐惧等，并给出每种情感的概率分数。

模型特点

高精度情感识别

能够准确识别多种语音情感状态，如悲伤、恐惧等。

基于Hubert架构

利用Hubert模型的强大特征提取能力进行情感分类。

概率输出

提供每种情感的概率分数，而不仅仅是单一分类结果。

模型能力

语音情感识别

音频分类

概率评分输出

使用案例

心理健康

情绪状态监测

通过语音分析监测用户的情绪变化

可识别悲伤、恐惧等负面情绪

人机交互

情感化语音助手

使语音助手能够根据用户情绪调整响应方式

提升交互体验

🚀 基于预训练模型的音频情感预测项目

本项目提供了一个使用预训练模型对本地音频文件进行情感预测的工作示例，借助HUBert模型实现音频分类，能够有效识别音频中的情感信息。

🚀 快速开始

以下是使用预训练模型预测本地音频文件情感的示例代码：

def predict_emotion_hubert(audio_file):
    """ inspired by an example from https://github.com/m3hrdadfi/soxan """
    from audio_models import HubertForSpeechClassification
    from transformers import  Wav2Vec2FeatureExtractor, AutoConfig
    import torch.nn.functional as F
    import torch
    import numpy as np
    from pydub import AudioSegment

    model = HubertForSpeechClassification.from_pretrained("Rajaram1996/Hubert_emotion") # Downloading: 362M
    feature_extractor = Wav2Vec2FeatureExtractor.from_pretrained("facebook/hubert-base-ls960")
    sampling_rate=16000 # defined by the model; must convert mp3 to this rate.
    config = AutoConfig.from_pretrained("Rajaram1996/Hubert_emotion")

    def speech_file_to_array(path, sampling_rate):
        # using torchaudio...
        # speech_array, _sampling_rate = torchaudio.load(path)
        # resampler = torchaudio.transforms.Resample(_sampling_rate, sampling_rate)
        # speech = resampler(speech_array).squeeze().numpy()
        sound = AudioSegment.from_file(path)
        sound = sound.set_frame_rate(sampling_rate)
        sound_array = np.array(sound.get_array_of_samples())
        return sound_array

    sound_array = speech_file_to_array(audio_file, sampling_rate)
    inputs = feature_extractor(sound_array, sampling_rate=sampling_rate, return_tensors="pt", padding=True)
    inputs = {key: inputs[key].to("cpu").float() for key in inputs}

    with torch.no_grad():
        logits = model(**inputs).logits

    scores = F.softmax(logits, dim=1).detach().cpu().numpy()[0]
    outputs = [{
        "emo": config.id2label[i],
        "score": round(score * 100, 1)}
        for i, score in enumerate(scores)
    ]
    return [row for row in sorted(outputs, key=lambda x:x["score"], reverse=True) if row['score'] != '0.0%'][:2]

基础用法

result = predict_emotion_hubert("male-crying.mp3")
>>> result
[{'emo': 'male_sad', 'score': 91.0}, {'emo': 'male_fear', 'score': 4.8}]

💡 使用建议

确保音频文件的采样率转换为模型所要求的16000Hz。

代码运行环境需安装audio_models、transformers、torch、numpy、pydub等相关依赖库。

📦 安装指南

由于原文档未提供具体的安装步骤，此部分暂不展示。

🔧 技术细节

由于原文档未提供具体的技术实现细节，此部分暂不展示。

📄 许可证

由于原文档未提供许可证信息，此部分暂不展示。

精选推荐AI模型

Llama 3 Typhoon V1.5x 8b Instruct

专为泰语设计的80亿参数指令模型，性能媲美GPT-3.5-turbo，优化了应用场景、检索增强生成、受限生成和推理任务

Cadet-Tiny是一个基于SODA数据集训练的超小型对话模型，专为边缘设备推理设计，体积仅为Cosmo-3B模型的2%左右。

Roberta Base Chinese Extractive Qa

基于RoBERTa架构的中文抽取式问答模型，适用于从给定文本中提取答案的任务。

智启未来，您的人工智能解决方案智库