English

A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion

Computer Vision and Pattern Recognition 2025-02-13 v1 Artificial Intelligence

Abstract

With the advancement of artificial intelligence and computer vision technologies, multimodal emotion recognition has become a prominent research topic. However, existing methods face challenges such as heterogeneous data fusion and the effective utilization of modality correlations. This paper proposes a novel multimodal emotion recognition approach, DeepMSI-MER, based on the integration of contrastive learning and visual sequence compression. The proposed method enhances cross-modal feature fusion through contrastive learning and reduces redundancy in the visual modality by leveraging visual sequence compression. Experimental results on two public datasets, IEMOCAP and MELD, demonstrate that DeepMSI-MER significantly improves the accuracy and robustness of emotion recognition, validating the effectiveness of multimodal feature fusion and the proposed approach.

Keywords

Cite

@article{arxiv.2502.08573,
  title  = {A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion},
  author = {Wei Dai and Dequan Zheng and Feng Yu and Yanrong Zhang and Yaohui Hou},
  journal= {arXiv preprint arXiv:2502.08573},
  year   = {2025}
}