English
Related papers

Related papers: To Fuse or to Drop? Dual-Path Learning for Resolvi…

200 papers

With the continuous development of deep learning (DL), the task of multimodal dialogue emotion recognition (MDER) has recently received extensive research attention, which is also an essential branch of DL. The MDER aims to identify the…

Computation and Language · Computer Science 2024-09-04 Wei Ai , Yuntao Shou , Tao Meng , Nan Yin , Keqin Li

Multimodal affective computing aims to predict humans' sentiment, emotion, intention, and opinion using language, acoustic, and visual modalities. However, current models often learn spurious correlations that harm generalization under…

Machine Learning · Computer Science 2026-04-21 Sijie Mai , Shiqin Han

Depression is a severe mental disorder, and reliable identification plays a critical role in early intervention and treatment. Multimodal depression detection aims to improve diagnostic performance by jointly modeling complementary…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Chongxiao Wang , Junjie Liang , Peng Cao , Jinzhu Yang , Osmar R. Zaiane

Multi-modality fusion and multi-task learning are becoming trendy in 3D autonomous driving scenario, considering robust prediction and computation budget. However, naively extending the existing framework to the domain of multi-modality…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Zhijian Huang , Sihao Lin , Guiyu Liu , Mukun Luo , Chaoqiang Ye , Hang Xu , Xiaojun Chang , Xiaodan Liang

Electroencephalogram (EEG) signals serve as a powerful tool in affective Brain-Computer Interfaces (aBCIs) and play a crucial role in affective computing. In recent years, the introduction of deep learning techniques has significantly…

Machine Learning · Computer Science 2025-08-08 Guangli Li , Canbiao Wu , Zhehao Zhou , Tuo Sun , Ping Tan , Li Zhang , Zhen Liang

Though multimodal emotion recognition has achieved significant progress over recent years, the potential of rich synergic relationships across the modalities is not fully exploited. In this paper, we introduce Recursive Joint Cross-Modal…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 R. Gnana Praveen , Jahangir Alam

Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional…

Robotics · Computer Science 2024-09-19 Youssef Mohamed , Severin Lemaignan , Arzu Guneysu , Patric Jensfelt , Christian Smith

In the domain of human-computer interaction, accurately recognizing and interpreting human emotions is crucial yet challenging due to the complexity and subtlety of emotional expressions. This study explores the potential for detecting a…

Multimedia · Computer Science 2025-05-13 Jiehui Jia , Huan Zhang , Jinhua Liang

With the rapid rise of social media and Internet culture, memes have become a popular medium for expressing emotional tendencies. This has sparked growing interest in Meme Emotion Understanding (MEU), which aims to classify the emotional…

Computation and Language · Computer Science 2025-11-17 Yi Shi , Wenlong Meng , Zhenyuan Guo , Chengkun Wei , Wenzhi Chen

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an innovative feature fusion framework based on cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jifeng Shen , Haibo Zhan , Xin Zuo , Heng Fan , Xiaohui Yuan , Jun Li , Wankou Yang

Current research on cross-modal retrieval is mostly English-oriented, as the availability of a large number of English-oriented human-labeled vision-language corpora. In order to break the limit of non-English labeled data, cross-lingual…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Yabing Wang , Shuhui Wang , Hao Luo , Jianfeng Dong , Fan Wang , Meng Han , Xun Wang , Meng Wang

Humans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of…

Machine Learning · Computer Science 2022-02-21 Vandana Rajan , Alessio Brutti , Andrea Cavallaro

Depression is a leading cause of death worldwide, and the diagnosis of depression is nontrivial. Multimodal learning is a popular solution for automatic diagnosis of depression, and the existing works suffer two main drawbacks: 1) the…

Multimedia · Computer Science 2023-01-03 Chengbo Yuan , Qianhui Xu , Yong Luo

Multimodal emotion recognition in conversations (MERC) requires integrating multimodal signals while being robust to noise and modeling contextual reasoning. Existing approaches often emphasize fusion but overlook uncertainty in noisy…

Computation and Language · Computer Science 2026-04-03 Yiqiang Cai , Chengyan Wu , Bolei Ma , Bo Chen , Yun Xue , Julia Hirschberg , Ziwei Gong

Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrates modalities such as…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Wenping Jin , Li Zhu , Jing Sun

Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker…

Sound · Computer Science 2025-11-19 Xiao Li , Kotaro Funakoshi , Manabu Okumura

Speech emotion recognition is a challenge and an important step towards more natural human-computer interaction (HCI). The popular approach is multimodal emotion recognition based on model-level fusion, which means that the multimodal…

Sound · Computer Science 2022-11-22 Fan Qian , Jiqing Han

Learning joint representations across multiple modalities remains a central challenge in multimodal machine learning. Prevailing approaches predominantly operate in pairwise settings, aligning two modalities at a time. While some recent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Stefanos Koutoupis , Michaela Areti Zervou , Konstantinos Kontras , Maarten De Vos , Panagiotis Tsakalides , Grigorios Tsagkatakis

Super-resolution (SR) is a severely ill-posed problem with inherent ambiguity, as widely recognized in both empirical and theoretical studies. Although recent semantic-guided and multi-modal SR methods exploit large models or external…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Jinyi Luo , Minghao Liu , Yifan Li , Zejia Fan , Jiaying Liu

We present a method for gesture detection and localisation based on multi-scale and multi-modal deep learning. Each visual modality captures spatial information at a particular spatial scale (such as motion of the upper body or a hand), and…

Computer Vision and Pattern Recognition · Computer Science 2015-07-21 Natalia Neverova , Christian Wolf , Graham W. Taylor , Florian Nebout