中文

情感语音消息 (EMOVOME) 数据库:自发语音消息中的情感识别

声音 2024-06-14 v2 人工智能 计算与语言 音频与语音处理

摘要

情感语音消息 (EMOVOME) 是一个自发语音数据集,包含来自 100 位性别平衡的西班牙语使用者在即时通讯应用真实对话中的 999 条音频消息。语音消息是在招募参与者之前于自然条件下产生的,避免了实验室环境带来的任何意识偏差。音频由三位非专家和两位专家在效价和唤醒度维度上进行标注,随后合并以获得每个维度的最终标签。专家还提供了对应于七种情感类别的额外标签。为使用 EMOVOME 的未来研究建立基线,我们利用语音和音频转录实现了情感识别模型。对于语音,我们使用了标准的 eGeMAPS 特征集和支持向量机,分别获得了 49.27% 和 44.71% 的未加权准确率(效价和唤醒度)。对于文本,我们微调了多语言 BERT 模型,分别实现了 61.15% 和 47.43% 的未加权准确率(效价和唤醒度)。该数据库将显著促进野外情感识别研究,同时为西班牙语提供独特、自然且免费可用的资源。

关键词

引用

@article{arxiv.2402.17496,
  title  = {Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages},
  author = {Lucía Gómez Zaragozá and Rocío del Amor and Elena Parra Vargas and Valery Naranjo and Mariano Alcañiz Raya and Javier Marín-Morales},
  journal= {arXiv preprint arXiv:2402.17496},
  year   = {2024}
}

备注

This paper has been superseded by arXiv:2403.02167 (merged from the description of the EMOVOME database in arXiv:2402.17496v1 and the speech emotion recognition models in arXiv:2403.02167v1)