English

MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition

Computer Vision and Pattern Recognition 2021-11-02 v1 Image and Video Processing

Abstract

Multimodal emotion recognition study is hindered by the lack of labelled corpora in terms of scale and diversity, due to the high annotation cost and label ambiguity. In this paper, we propose a pre-training model \textbf{MEmoBERT} for multimodal emotion recognition, which learns multimodal joint representations through self-supervised learning from large-scale unlabeled video data that come in sheer volume. Furthermore, unlike the conventional "pre-train, finetune" paradigm, we propose a prompt-based method that reformulates the downstream emotion classification task as a masked text prediction one, bringing the downstream task closer to the pre-training. Extensive experiments on two benchmark datasets, IEMOCAP and MSP-IMPROV, show that our proposed MEmoBERT significantly enhances emotion recognition performance.

Keywords

Cite

@article{arxiv.2111.00865,
  title  = {MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition},
  author = {Jinming Zhao and Ruichen Li and Qin Jin and Xinchao Wang and Haizhou Li},
  journal= {arXiv preprint arXiv:2111.00865},
  year   = {2021}
}

Comments

4 papges, 2 figures

R2 v1 2026-06-24T07:20:44.314Z