相似多于相异:面向跨语料库语音情感识别的标注者建模
声音
2025-09-17 v1 机器学习
音频与语音处理
摘要
语音情感识别系统通常预测由多名标注者评分生成的共识值。然而,这些模型预测任意单个人标注的能力有限。或者,模型可以学习预测所有标注者的标注。将此类模型适应于新标注者十分困难,因为新标注者必须单独提供足够的标注训练数据。我们提出利用标注者间的相似性,通过使用在大量标注者群体上预训练的模型来识别相似的、先前见过的标注者。给定一个新的、未见过的标注者和有限的注册数据,我们可以为相似的标注者进行预测,从而实现对目标数据集中未见数据的开箱即用标注,提供一种极低成本的个性化机制。我们证明我们的方法显著优于其他开箱即用的方法,为实际部署中可行的轻量级情感适应铺平了道路。
引用
@article{arxiv.2509.12295,
title = {More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition},
author = {James Tavernor and Emily Mower Provost},
journal= {arXiv preprint arXiv:2509.12295},
year = {2025}
}
备注
\copyright 20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works