English

Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition

Sound 2026-02-24 v1

Abstract

This paper highlights the critical importance of multi-channel speech enhancement (MCSE) for speech emotion recognition (ER) in cocktail party scenarios. A multi-channel speech dereverberation and separation front-end integrating DNN-WPE and mask-based MVDR is used to extract the target speaker's speech from the mixture speech, before being fed into the downstream ER back-end using HuBERT- and ViT-based speech and visual features. Experiments on mixture speech constructed using the IEMOCAP and MSP-FACE datasets suggest the MCSE output consistently outperforms domain fine-tuned single-channel speech representations produced by: a) Conformer-based metric GANs; and b) WavLM SSL features with optional SE-ER dual task fine-tuning. Statistically significant increases in weighted, unweighted accuracy and F1 measures by up to 9.5%, 8.5% and 9.1% absolute (17.1%, 14.7% and 16.0% relative) are obtained over the above single-channel baselines. The generalization of IEMOCAP trained MCSE front-ends are also shown when being zero-shot applied to out-of-domain MSP-FACE data.

Keywords

Cite

@article{arxiv.2602.18802,
  title  = {Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition},
  author = {Youjun Chen and Guinan Li and Mengzhe Geng and Xurong Xie and Shujie Hu and Huimeng Wang and Haoning Xu and Chengxi Deng and Jiajun Deng and Zhaoqing Li and Mingyu Cui and Xunying Liu},
  journal= {arXiv preprint arXiv:2602.18802},
  year   = {2026}
}

Comments

Accepted by ICASSP2026