English

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

Multimedia 2026-07-19 v1 Computation and Language Machine Learning

Abstract

Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across utterances caused by conflicting cues, varying noise, and missing modality-specific signals. We propose EmoEUS, an explicit uncertainty supervision framework for MERC. EmoEUS performs uncertainty-aware multimodal fusion by dynamically weighting modalities using learned variance estimates. We also introduce an explicitly supervised loss that aligns each utterance's predicted variance with the distance between the utterance's distributional representation and its emotion- and modality-specific cluster center. Experiments on IEMOCAP and MELD show that EmoEUS consistently outperforms state-of-the-art methods.

Keywords

Cite

@article{arxiv.2607.18336,
  title  = {EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation},
  author = {Zilong Huang and Kong Aik Lee and Junjie Li and Zhe Li and Man-Wai Mak},
  journal= {arXiv preprint arXiv:2607.18336},
  year   = {2026}
}

Comments

Accept by Interspeech 2026