English
Related papers

Related papers: MAVEN: Multi-modal Attention for Valence-Arousal E…

200 papers

In this paper, we consider the problem of real-time video-based facial emotion analytics, namely, facial expression recognition, prediction of valence and arousal and detection of action unit points. We propose the novel frame-level emotion…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Andrey V. Savchenko

For a long time, images have proved perfect at both storing and conveying rich semantics, especially human emotions. A lot of research has been conducted to provide machines with the ability to recognize emotions in photos of people.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Quoc-Bao Ninh , Hai-Chan Nguyen , Triet Huynh , Trung-Nghia Le

Over the past few years many research efforts have been devoted to the field of affect analysis. Various approaches have been proposed for: i) discrete emotion recognition in terms of the primary facial expressions; ii) emotion analysis in…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Dimitrios Kollias , Stefanos Zafeiriou

Human emotion recognition holds a pivotal role in facilitating seamless human-computer interaction. This paper delineates our methodology in tackling the Valence-Arousal (VA) Estimation Challenge, Expression (Expr) Classification Challenge,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Weiwei Zhou , Jiada Lu , Chenkun Ling , Weifeng Wang , Shaowei Liu

This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jun Yu , Gongpeng Zhao , Yongqi Wang , Zhihong Wei , Yang Zheng , Zerui Zhang , Zhongpeng Cai , Guochen Xie , Jichao Zhu , Wangyuan Zhu

Predicting the emotional impact of videos using machine learning is a challenging task considering the varieties of modalities, the complicated temporal contex of the video as well as the time dependency of the emotional states. Feature…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Jie Zhang , Yin Zhao , Longjun Cai , Chaoping Tu , Wu Wei

Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Jun Yu , Naixiang Zheng , Guoyuan Wang , Yunxiang Zhang , Lingsi Zhu , Jiaen Liang , Wei Huang , Shengping Liu

Emotion Recognition in Conversations (ERC) is crucial in developing sympathetic human-machine interaction. In conversational videos, emotion can be present in multiple modalities, i.e., audio, video, and transcript. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Vishal Chudasama , Purbayan Kar , Ashish Gudmalwar , Nirmesh Shah , Pankaj Wasnik , Naoyuki Onoe

To address the limitation in multimodal emotion recognition (MER) performance arising from inter-modal information fusion, we propose a novel MER framework based on multitask learning where fusion occurs after alignment, called Foal-Net.…

Multimedia · Computer Science 2024-08-20 Qifei Li , Yingming Gao , Yuhua Wen , Cong Wang , Ya Li

Detecting emotional inconsistency across modalities is a key challenge in affective computing, as speech and text often convey conflicting cues. Existing approaches generally rely on incomplete emotion representations and employ…

Multimedia · Computer Science 2025-09-25 Zongyi Li , Junchuan Zhao , Francis Bu Sung Lee , Andrew Zi Han Yee

Understanding dynamic scenes and dialogue contexts in order to converse with users has been challenging for multimodal dialogue systems. The 8-th Dialog System Technology Challenge (DSTC8) proposed an Audio Visual Scene-Aware Dialog (AVSD)…

Computation and Language · Computer Science 2020-01-20 Yun-Wei Chu , Kuan-Yen Lin , Chao-Chun Hsu , Lun-Wei Ku

Video affective understanding, which aims to predict the evoked expressions by the video content, is desired for video creation and recommendation. In the recent EEV challenge, a dense affective understanding task is proposed and requires…

Computer Vision and Pattern Recognition · Computer Science 2021-06-21 Baoming Yan , Lin Wang , Ke Gao , Bo Gao , Xiao Liu , Chao Ban , Jiang Yang , Xiaobo Li

Multimodal emotion recognition plays a key role in many domains, including mental health monitoring, educational interaction, and human-computer interaction. However, existing methods often face three major challenges: unbalanced category…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Feng Li , Ke Wu , Yongwei Li

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Han Zhang , Wanting Jiang , Tomasz Kornuta , Tian Zheng , Vidya Murali

Emotion recognition in user-generated videos plays an important role in human-centered computing. Existing methods mainly employ traditional two-stage shallow pipeline, i.e. extracting visual and/or audio features and training classifiers.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-03 Sicheng Zhao , Yunsheng Ma , Yang Gu , Jufeng Yang , Tengfei Xing , Pengfei Xu , Runbo Hu , Hua Chai , Kurt Keutzer

Applications of an efficient emotion recognition system can be found in several domains such as medicine, driver fatigue surveillance, social robotics, and human-computer interaction. Appraising human emotional states, behaviors, and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Savinay Nagendra , Prapti Panigrahi

We propose a cross-modal co-attention model for continuous emotion recognition using visual-audio-linguistic information. The model consists of four blocks. The visual, audio, and linguistic blocks are used to learn the spatial-temporal…

Multimedia · Computer Science 2022-03-31 Su Zhang , Ruyi An , Yi Ding , Cuntai Guan

Current FER (Facial Expression Recognition) dataset is mostly labeled by emotion categories, such as happy, angry, sad, fear, disgust, surprise, and neutral which are limited in expressiveness. However, future affective computing requires…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yi Huo , Yun Ge

We propose an audio-visual spatial-temporal deep neural network with: (1) a visual block containing a pretrained 2D-CNN followed by a temporal convolutional network (TCN); (2) an aural block containing several parallel TCNs; and (3) a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Su Zhang , Yi Ding , Ziquan Wei , Cuntai Guan

Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Junhyeong Byeon , Jeongyeol Kim , Sejoon Lim