English

Convolutional Attention Networks for Multimodal Emotion Recognition from Speech and Text Data

Computation and Language 2019-03-11 v2 Artificial Intelligence Human-Computer Interaction

Abstract

Emotion recognition has become a popular topic of interest, especially in the field of human computer interaction. Previous works involve unimodal analysis of emotion, while recent efforts focus on multi-modal emotion recognition from vision and speech. In this paper, we propose a new method of learning about the hidden representations between just speech and text data using convolutional attention networks. Compared to the shallow model which employs simple concatenation of feature vectors, the proposed attention model performs much better in classifying emotion from speech and text data contained in the CMU-MOSEI dataset.

Keywords

Cite

@article{arxiv.1805.06606,
  title  = {Convolutional Attention Networks for Multimodal Emotion Recognition from Speech and Text Data},
  author = {Chan Woo Lee and Kyu Ye Song and Jihoon Jeong and Woo Yong Choi},
  journal= {arXiv preprint arXiv:1805.06606},
  year   = {2019}
}

Comments

Inaccurate scientific facts listed in document

R2 v1 2026-06-23T01:58:18.814Z