词汇特征在语音情感识别中的贡献
音频与语音处理
2025-09-09 v1 计算与语言
声音
摘要
尽管在语音情感识别 (SER) 中常认为语音的副语言线索是主要驱动因素,但我们调查从语音中提取的词汇内容的作用,发现其可实现与声学模型相当甚至更好的性能。在 MELD 数据集上,我们的基于词汇的方法获得加权 F1 分数 (WF1) 为 51.5%,而声学-only 管道(参数更大)仅为 49.3%。此外,我们分析了不同的自监督学习 (SSL) 语音和文本表示,进行基于 transformer 编码器的层级研究,并评估了音频去噪的效果。
引用
@article{arxiv.2509.05634,
title = {On the Contribution of Lexical Features to Speech Emotion Recognition},
author = {David Combei},
journal= {arXiv preprint arXiv:2509.05634},
year = {2025}
}
备注
Accepted to 13th Conference on Speech Technology and Human-Computer Dialogue