English
Related papers

Related papers: Mamba-Enhanced Text-Audio-Video Alignment Network …

200 papers

Emotion recognition in conversation (ERC) has attracted much attention in recent years for its necessity in widespread applications. Existing ERC methods mostly model the self and inter-speaker context separately, posing a major issue for…

Computation and Language · Computer Science 2021-12-24 Chen Liang , Chong Yang , Jing Xu , Juyang Huang , Yongliang Wang , Yang Dong

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

Recognizing the emotional state of people is a basic but challenging task in video understanding. In this paper, we propose a new task in this field, named Pairwise Emotional Relationship Recognition (PERR). This task aims to recognize the…

Computer Vision and Pattern Recognition · Computer Science 2021-09-24 Xun Gao , Yin Zhao , Jie Zhang , Longjun Cai

Despite the recent progress in speech emotion recognition (SER), state-of-the-art systems are unable to achieve improved performance in cross-language settings. In this paper, we propose a Multimodal Dual Attention Transformer (MDAT) model…

Computation and Language · Computer Science 2023-07-17 Syed Aun Muhammad Zaidi , Siddique Latif , Junaid Qadir

Emotion recognition from EEG signals is essential for affective computing and has been widely explored using deep learning. While recent deep learning approaches have achieved strong performance on single EEG emotion datasets, their…

Machine Learning · Computer Science 2025-11-17 Yuning Chen , Sha Zhao , Shijian Li , Gang Pan

To establish empathy with machines, it is essential to fully understand human emotional changes. However, research in multimodal emotion recognition often overlooks one problem: individual expressive traits vary significantly, which means…

Sound · Computer Science 2026-04-29 Kexue Wang , Yinfeng Yu , Liejun Wang

Multimodal Sentiment Analysis (MSA) aims to recognize human emotions by exploiting textual, acoustic, and visual modalities, and thus how to make full use of the interactions between different modalities is a central challenge of MSA.…

Computation and Language · Computer Science 2025-02-17 Yubo Gao , Haotian Wu , Lei Zhang

Multimodal Emotion Recognition refers to the classification of input video sequences into emotion labels based on multiple input modalities (usually video, audio and text). In recent years, Deep Neural networks have shown remarkable…

Machine Learning · Computer Science 2024-10-28 Ashish Ramayee Asokan , Nidarshan Kumar , Anirudh Venkata Ragam , Shylaja S Sharath

Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only one modality is available and can be implemented…

Machine Learning · Computer Science 2026-01-13 Lucas Goncalves , Seong-Gyun Leem , Wei-Cheng Lin , Berrak Sisman , Carlos Busso

Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in applications such as social robots and human-computer interaction, where understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 He Hu , Tengjin Weng , Zebang Cheng , Yu Wang , Jiachen Luo , Björn Schuller , Zheng Lian , Laizhong Cui

Multimodal emotion recognition (MER) extracts emotions from multimodal data, including visual, speech, and text inputs, playing a key role in human-computer interaction. Attention-based fusion methods dominate MER research, achieving strong…

Artificial Intelligence · Computer Science 2025-06-03 Jiajun He , Jinyi Mi , Tomoki Toda

Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Kaisi Guan , Zhengfeng Lai , Yuchong Sun , Peng Zhang , Wei Liu , Kieran Liu , Meng Cao , Ruihua Song

With the extensive accumulation of conversational data on the Internet, emotion recognition in conversations (ERC) has received increasing attention. Previous efforts of this task mainly focus on leveraging contextual and speaker-specific…

Computation and Language · Computer Science 2023-06-28 Yinyi Wei , Shuaipeng Liu , Hailei Yan , Wei Ye , Tong Mo , Guanglu Wan

Multimodal Emotion Recognition in Conversation (MERC) aims to predict speakers' emotions by integrating textual, acoustic, and visual cues. Existing approaches either struggle to capture complex cross-modal interactions or experience…

Multimedia · Computer Science 2026-03-24 Xiaosen Lyu , Jiayu Xiong , Yuren Chen , Wanlong Wang , Xiaoqing Dai , Jing Wang

Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person of interest conditioned on natural language descriptions. Distinct from conventional action recognition and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Kunyu Peng , Di Wen , Jia Fu , Jiamin Wu , Kailun Yang , Junwei Zheng , Ruiping Liu , Yufan Chen , Yuqian Fu , Danda Pani Paudel , Luc Van Gool , Rainer Stiefelhagen

The prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker's emotion, with the text generally obtained through automatic speech recognition…

Computation and Language · Computer Science 2024-05-29 Jiajun He , Xiaohan Shi , Xingfeng Li , Tomoki Toda

The lack of data and the difficulty of multimodal fusion have always been challenges for multimodal emotion recognition (MER). In this paper, we propose to use pretrained models as upstream network, wav2vec 2.0 for audio modality and BERT…

Computation and Language · Computer Science 2023-02-28 Dekai Sun , Yancheng He , Jiqing Han

In this paper, a novel two-branch neural network model structure is proposed for multimodal emotion recognition, which consists of a time synchronous branch (TSB) and a time asynchronous branch (TAB). To capture correlations between each…

Computation and Language · Computer Science 2021-07-23 Wen Wu , Chao Zhang , Philip C. Woodland

Multimodal Emotion Recognition in Conversation (MERC) significantly enhances emotion recognition performance by integrating complementary emotional cues from text, audio, and visual modalities. While existing methods commonly utilize…

Multimedia · Computer Science 2026-02-12 Xinyi Che , Wenbo Wang , Jian Guan , Qijun Zhao

Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities are inconsistent.…

Artificial Intelligence · Computer Science 2025-10-14 Zhiyuan Han , Beier Zhu , Yanlong Xu , Peipei Song , Xun Yang
‹ Prev 1 8 9 10 Next ›