English

Group Gated Fusion on Attention-based Bidirectional Alignment for Multimodal Emotion Recognition

Computation and Language 2022-01-19 v1 Sound Audio and Speech Processing

Abstract

Emotion recognition is a challenging and actively-studied research area that plays a critical role in emotion-aware human-computer interaction systems. In a multimodal setting, temporal alignment between different modalities has not been well investigated yet. This paper presents a new model named as Gated Bidirectional Alignment Network (GBAN), which consists of an attention-based bidirectional alignment network over LSTM hidden states to explicitly capture the alignment relationship between speech and text, and a novel group gated fusion (GGF) layer to integrate the representations of different modalities. We empirically show that the attention-aligned representations outperform the last-hidden-states of LSTM significantly, and the proposed GBAN model outperforms existing state-of-the-art multimodal approaches on the IEMOCAP dataset.

Keywords

Cite

@article{arxiv.2201.06309,
  title  = {Group Gated Fusion on Attention-based Bidirectional Alignment for Multimodal Emotion Recognition},
  author = {Pengfei Liu and Kun Li and Helen Meng},
  journal= {arXiv preprint arXiv:2201.06309},
  year   = {2022}
}

Comments

Published in INTERSPEECH-2020

R2 v1 2026-06-24T08:52:08.292Z