English

deep learning of segment-level feature representation for speech emotion recognition in conversations

Computation and Language 2023-02-07 v1 Sound Audio and Speech Processing

Abstract

Accurately detecting emotions in conversation is a necessary yet challenging task due to the complexity of emotions and dynamics in dialogues. The emotional state of a speaker can be influenced by many different factors, such as interlocutor stimulus, dialogue scene, and topic. In this work, we propose a conversational speech emotion recognition method to deal with capturing attentive contextual dependency and speaker-sensitive interactions. First, we use a pretrained VGGish model to extract segment-based audio representation in individual utterances. Second, an attentive bi-directional gated recurrent unit (GRU) models contextual-sensitive information and explores intra- and inter-speaker dependencies jointly in a dynamic manner. The experiments conducted on the standard conversational dataset MELD demonstrate the effectiveness of the proposed method when compared against state-of the-art methods.

Keywords

Cite

@article{arxiv.2302.02419,
  title  = {deep learning of segment-level feature representation for speech emotion recognition in conversations},
  author = {Jiachen Luo and Huy Phan and Joshua Reiss},
  journal= {arXiv preprint arXiv:2302.02419},
  year   = {2023}
}

Comments

6 pages, 4 figures

R2 v1 2026-06-28T08:32:25.166Z