中文
相关论文

相关论文: MicroEmo: Time-Sensitive Multimodal Emotion Recogn…

200 篇论文

Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Hao Cheng , Zhiwei Zhao , Yichao He , Zhenzhen Hu , Jia Li , Meng Wang , Richang Hong

This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user studies report an…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tianwei Chen , Takuya Furusawa , Yuki Hirakawa , Ryotaro Shimizu , Mo Fan , Takashi Wada

Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demonstrate robust general capabilities, they face considerable…

计算与语言 · 计算机科学 2025-07-08 Ao Li , Longwei Xu , Chen Ling , Jinghui Zhang , Pengwei Wang

Emotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which raises concerns about personal privacy. To address this, we…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Deng Li , Bohao Xing , Xin Liu , Baiqiang Xia , Bihan Wen , Heikki Kälviäinen

Facial expression is temporally dynamic event which can be decomposed into a set of muscle motions occurring in different facial regions over various time intervals. For dynamic expression recognition, two key issues, temporal alignment and…

计算机视觉与模式识别 · 计算机科学 2016-11-23 Mengyi Liu , Shiguang Shan , Ruiping Wang , Xilin Chen

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilities, they tend to…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Yuxiang Guo , Faizan Siddiqui , Yang Zhao , Rama Chellappa , Shao-Yuan Lo

Emotion recognition is a core research area at the intersection of artificial intelligence and human communication analysis. It is a significant technical challenge since humans display their emotions through complex idiosyncratic…

人机交互 · 计算机科学 2018-09-14 Paul Pu Liang , Amir Zadeh , Louis-Philippe Morency

Spatial-temporal feature learning is of vital importance for video emotion recognition. Previous deep network structures often focused on macro-motion which extends over long time scales, e.g., on the order of seconds. We believe…

计算机视觉与模式识别 · 计算机科学 2019-11-25 Didan Deng , Zhaokang Chen , Yuqian Zhou , Bertram Shi

Understanding and predicting emotion from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis,…

计算机视觉与模式识别 · 计算机科学 2025-11-05 Zhicheng Zhang , Weicheng Wang , Yongjie Zhu , Wenyu Qin , Pengfei Wan , Di Zhang , Jufeng Yang

As a spontaneous expression of emotion on face, micro-expression reveals the underlying emotion that cannot be controlled by human. In micro-expression, facial movement is transient and sparsely localized through time. However, the existing…

计算机视觉与模式识别 · 计算机科学 2021-03-16 Min Peng , Chongyang Wang , Yuan Gao , Tao Bi , Tong Chen , Yu Shi , Xiang-Dong Zhou

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current…

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

Multi-modal conversation emotion recognition (MCER) aims to recognize and track the speaker's emotional state using text, speech, and visual information in the conversation scene. Analyzing and studying MCER issues is significant to…

人工智能 · 计算机科学 2025-11-14 Yuntao Shou , Tao Meng , Wei Ai , Fangze Fu , Nan Yin , Keqin Li

Recognising emotions in context involves identifying an individual's apparent emotions while considering contextual cues from the surrounding scene. Previous approaches to this task have typically designed explicit scene-encoding…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Alexandros Xenos , Niki Maria Foteinopoulou , Ioanna Ntinou , Ioannis Patras , Georgios Tzimiropoulos

Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made tremendous progress in the last few years for many visual tasks, potentially offering a…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Madhav Agarwal , Sotirios A. Tsaftaris , Laura Sevilla-Lara , Steven McDonagh

The integration of information across multiple modalities and across time is a promising way to enhance the emotion recognition performance of affective systems. Much previous work has focused on instantaneous emotion recognition. The 2018…

图像与视频处理 · 电气工程与系统科学 2018-05-07 Didan Deng , Yuqian Zhou , Jimin Pi , Bertram E. Shi

Multimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper understanding of human…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Jinpeng Hu , Hongchang Shi , Chongyuan Dai , Zhuo Li , Peipei Song , Meng Wang

Micro-expression recognition plays a pivotal role in understanding hidden emotions and has applications across various fields. Traditional recognition methods assume access to all training data at once, but real-world scenarios involve…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zhengqin Lai , Xiaopeng Hong , Yabin Wang , Xiaobai Li

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While Multimodal Large Language Models (MLLMs) are natural…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Liangyang Ouyang , Yifei Huang , Mingfang Zhang , Caixin Kang , Ryosuke Furuta , Yoichi Sato

The performance of speech emotion recognition (SER) is limited by the insufficient emotion information in unimodal systems and the feature alignment difficulties in multimodal systems. Recently, multimodal large language models (MLLMs) have…

声音 · 计算机科学 2025-09-22 Yiqing Yang , Man-Wai Mak