中文

元数据增强语音情感识别:增强残差集成与双阶段微调中的协注意力机制

音频与语音处理 2025-04-29 v1 声音

摘要

语音情感识别(SER)涉及分析人声表达以确定说话者的情绪状态,其中充分且彻底地利用音频信息至关重要。因此,我们提出了一种新方法,利用自监督学习(SSL)模型并充分利用所有可用的辅助信息—— Specifically 指的是元数据——来增强性能。通过多任务学习中的双阶段微调方法,我们引入了增强残差集成(ARI)模块,用于增强 SSL 模型编码器中的 Transformer 层。该模块有效地保留了来自不同层次的声学特征,从而显著提高了需要各种特征水平的元数据相关辅助任务的性能。此外,由于 ARI 与其互补性,我们将协注意力模块纳入考虑,使模型能够有效地利用来自元数据相关辅助任务的多维信息和上下文关系。在预训练基模型和说话者独立设置下,我们的方法在 IEMOCAP 数据集上 consistently 超过了最新(SOTA)模型。

关键词

引用

@article{arxiv.2412.20707,
  title  = {Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning},
  author = {Zixiang Wan and Ziyue Qiu and Yiyang Liu and Wei-Qiang Zhang},
  journal= {arXiv preprint arXiv:2412.20707},
  year   = {2025}
}

备注

accepted by ICASSP2025. \c{opyright}2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component