English

More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition

Sound 2025-09-17 v1 Machine Learning Audio and Speech Processing

Abstract

Speech emotion recognition systems often predict a consensus value generated from the ratings of multiple annotators. However, these models have limited ability to predict the annotation of any one person. Alternatively, models can learn to predict the annotations of all annotators. Adapting such models to new annotators is difficult as new annotators must individually provide sufficient labeled training data. We propose to leverage inter-annotator similarity by using a model pre-trained on a large annotator population to identify a similar, previously seen annotator. Given a new, previously unseen, annotator and limited enrollment data, we can make predictions for a similar annotator, enabling off-the-shelf annotation of unseen data in target datasets, providing a mechanism for extremely low-cost personalization. We demonstrate our approach significantly outperforms other off-the-shelf approaches, paving the way for lightweight emotion adaptation, practical for real-world deployment.

Keywords

Cite

@article{arxiv.2509.12295,
  title  = {More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition},
  author = {James Tavernor and Emily Mower Provost},
  journal= {arXiv preprint arXiv:2509.12295},
  year   = {2025}
}

Comments

\copyright 20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

R2 v1 2026-07-01T05:37:36.159Z