English
Related papers

Related papers: To Fuse or to Drop? Dual-Path Learning for Resolvi…

200 papers

Multimodal emotion recognition aims to recognize emotions for each utterance of multiple modalities, which has received increasing attention for its application in human-machine interaction. Current graph-based methods fail to…

Computation and Language · Computer Science 2023-11-21 Dongyuan Li , Yusong Wang , Kotaro Funakoshi , Manabu Okumura

Emotion recognition plays a vital role in enhancing human-computer interaction. In this study, we tackle the MER-SEMI challenge of the MER2025 competition by proposing a novel multimodal emotion recognition framework. To address the issue…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Juewen Hu , Yexin Li , Jiulin Li , Shuo Chen , Pring Wong

Dynamic Facial Expression Recognition (DFER) aims to identify human emotions from temporally evolving facial movements and plays a critical role in affective computing. While recent vision-language approaches have introduced semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Yu Liu , Leyuan Qu , Hanlei Shi , Di Gao , Yuhua Zheng , Taihao Li

Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Yong Li , Yuanzhi Wang , Yi Ding , Shiqing Zhang , Ke Lu , Cuntai Guan

Multimodal multi-label emotion recognition (MMER) aims to identify the concurrent presence of multiple emotions in multimodal data. Existing studies primarily focus on improving fusion strategies and modeling modality-to-label dependencies.…

Computation and Language · Computer Science 2025-02-20 Jingwang Huang , Jiang Zhong , Qin Lei , Jinpeng Gao , Yuming Yang , Sirui Wang , Peiguang Li , Kaiwen Wei

Recent advancements in Multimodal Emotion Recognition (MER) face challenges in addressing both modality missing and Out-Of-Distribution (OOD) data simultaneously. Existing methods often rely on specific models or introduce excessive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Guowei Zhong , Ruohong Huan , Mingzhen Wu , Ronghua Liang , Peng Chen

This paper presents a novel deep neural network (DNN) for multimodal fusion of audio, video and text modalities for emotion recognition. The proposed DNN architecture has independent and shared layers which aim to learn the representation…

Computer Vision and Pattern Recognition · Computer Science 2019-07-09 Juan D. S. Ortega , Mohammed Senoussaoui , Eric Granger , Marco Pedersoli , Patrick Cardinal , Alessandro L. Koerich

In terms of human-computer interaction, it is becoming more and more important to correctly understand the user's emotional state in a conversation, so the task of multimodal emotion recognition (MER) started to receive more attention.…

Computation and Language · Computer Science 2024-01-04 Wei Ai , FuChen Zhang , Tao Meng , YunTao Shou , HongEn Shao , Keqin Li

Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-modal reasoning remains underexplored, with conflicting reports on whether added modalities help or…

Computation and Language · Computer Science 2026-05-01 Yucheng Wang , Yifan Hou , Aydin Javadov , Mubashara Akhtar , Mrinmaya Sachan

Since Multimodal Emotion Recognition in Conversation (MERC) can be applied to public opinion monitoring, intelligent dialogue robots, and other fields, it has received extensive research attention in recent years. Unlike traditional…

Machine Learning · Computer Science 2024-07-25 Tao Meng , Fuchen Zhang , Yuntao Shou , Hongen Shao , Wei Ai , Keqin Li

Multimodal Emotion Recognition in Conversations (MERC) identifies emotional states across text, audio and video, which is essential for intelligent dialogue systems and opinion analysis. Existing methods emphasize heterogeneous modal fusion…

Machine Learning · Computer Science 2025-04-01 Jiagen Li , Rui Yu , Huihao Huang , Huaicheng Yan

Multimodal Sentiment Analysis (MSA) aims to recognize human emotions by exploiting textual, acoustic, and visual modalities, and thus how to make full use of the interactions between different modalities is a central challenge of MSA.…

Computation and Language · Computer Science 2025-02-17 Yubo Gao , Haotian Wu , Lei Zhang

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

Multimedia · Computer Science 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang

One of the most significant challenges in Music Emotion Recognition (MER) comes from the fact that emotion labels can be heterogeneous across datasets with regard to the emotion representation, including categorical (e.g., happy, sad)…

Sound · Computer Science 2025-04-14 Jaeyong Kang , Dorien Herremans

Cognitive diagnosis model (CDM) is a fundamental and upstream component in intelligent education. It aims to infer students' mastery levels based on historical response logs. However, existing CDMs usually follow the ID-based embedding…

Artificial Intelligence · Computer Science 2024-10-22 Yuanhao Liu , Shuo Liu , Yimeng Liu , Jingwen Yang , Hong Qian

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

Image-Text Retrieval (ITR) is challenging in bridging visual and lingual modalities. Contrastive learning has been adopted by most prior arts. Except for limited amount of negative image-text pairs, the capability of constrastive learning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Haoran Wang , Dongliang He , Wenhao Wu , Boyang Xia , Min Yang , Fu Li , Yunlong Yu , Zhong Ji , Errui Ding , Jingdong Wang

Ambivalence and hesitancy (A/H) are subtle affective states where a person shows conflicting signals through different channels -- saying one thing while their face or voice tells another story. Recognising these states automatically is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Salah Eddine Bekhouche , Hichem Telli , Azeddine Benlamoudi , Salah Eddine Herrouz , Abdelmalik Taleb-Ahmed , Abdenour Hadid

Multimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are particularly important…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-07 Jiachen Luo , Huy Phan , Joshua Reiss

Multimodal Emotion Recognition in Conversation (MERC) aims to enhance emotion understanding by integrating complementary cues from text, audio, and visual modalities. Existing MERC approaches predominantly focus on cross-modal shared…

Multimedia · Computer Science 2025-12-16 Xinyi Che , Wenbo Wang , Yuanbo Hou , Mingjie Xie , Qijun Zhao , Jian Guan