English

Video Multimodal Emotion Recognition System for Real World Applications

Human-Computer Interaction 2023-08-29 v1

Abstract

This paper proposes a system capable of recognizing a speaker's utterance-level emotion through multimodal cues in a video. The system seamlessly integrates multiple AI models to first extract and pre-process multimodal information from the raw video input. Next, an end-to-end MER model sequentially predicts the speaker's emotions at the utterance level. Additionally, users can interactively demonstrate the system through the implemented interface.

Keywords

Cite

@article{arxiv.2308.14320,
  title  = {Video Multimodal Emotion Recognition System for Real World Applications},
  author = {Sun-Kyung Lee and Jong-Hwan Kim},
  journal= {arXiv preprint arXiv:2308.14320},
  year   = {2023}
}

Comments

Interspeech 2023

R2 v1 2026-06-28T12:05:43.331Z