English

MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition

Audio and Speech Processing 2024-06-27 v3 Computation and Language Sound

Abstract

Speech Emotion Recognition (SER) is an important research topic in human-computer interaction. Many recent works focus on directly extracting emotional cues through pre-trained knowledge, frequently overlooking considerations of appropriateness and comprehensiveness. Therefore, we propose a novel framework for pre-training knowledge in SER, called Multi-perspective Fusion Search Network (MFSN). Considering comprehensiveness, we partition speech knowledge into Textual-related Emotional Content (TEC) and Speech-related Emotional Content (SEC), capturing cues from both semantic and acoustic perspectives, and we design a new architecture search space to fully leverage them. Considering appropriateness, we verify the efficacy of different modeling approaches in capturing SEC and fills the gap in current research. Experimental results on multiple datasets demonstrate the superiority of MFSN.

Keywords

Cite

@article{arxiv.2306.09361,
  title  = {MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition},
  author = {Haiyang Sun and Fulin Zhang and Yingying Gao and Zheng Lian and Shilei Zhang and Junlan Feng},
  journal= {arXiv preprint arXiv:2306.09361},
  year   = {2024}
}
R2 v1 2026-06-28T11:06:23.854Z