English

Learning Robust Self-attention Features for Speech Emotion Recognition with Label-adaptive Mixup

Computation and Language 2023-05-11 v1 Sound Audio and Speech Processing

Abstract

Speech Emotion Recognition (SER) is to recognize human emotions in a natural verbal interaction scenario with machines, which is considered as a challenging problem due to the ambiguous human emotions. Despite the recent progress in SER, state-of-the-art models struggle to achieve a satisfactory performance. We propose a self-attention based method with combined use of label-adaptive mixup and center loss. By adapting label probabilities in mixup and fitting center loss to the mixup training scheme, our proposed method achieves a superior performance to the state-of-the-art methods.

Keywords

Cite

@article{arxiv.2305.06273,
  title  = {Learning Robust Self-attention Features for Speech Emotion Recognition with Label-adaptive Mixup},
  author = {Lei Kang and Lichao Zhang and Dazhi Jiang},
  journal= {arXiv preprint arXiv:2305.06273},
  year   = {2023}
}

Comments

Accepted to ICASSP 2023

R2 v1 2026-06-28T10:31:15.577Z