English
Related papers

Related papers: OMG - Emotion Challenge Solution

200 papers

The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions for the Semi-Supervised Learning Sub-Challenge…

Sound · Computer Science 2024-09-10 Qi Fan , Yutong Li , Yi Xin , Xinyu Cheng , Guanglai Gao , Miao Ma

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand the…

Sound · Computer Science 2025-11-21 Suhwan Choi , Kyu Won Kim , Myungjoo Kang

With the increasing demands of emotion comprehension and regulation in our daily life, a customized music-based emotion regulation system is introduced by employing current EEG information and song features, which predicts users' emotion…

Human-Computer Interaction · Computer Science 2022-11-29 Jiyang Li , Wei Wang , Kratika Bhagtani , Yincheng Jin , Zhanpeng Jin

The main idea of this ISO is to use StarGAN (A type of GAN model) to perform training and testing on an emotion dataset resulting in a emotion recognition which can be generated by the valence arousal score of the 7 basic expressions. We…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Aritra Banerjee , Dimitrios Kollias

This paper details the sixth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW 2018 is a grand challenge in the ACM International Conference on Multimodal Interaction 2018, Colorado, USA. The challenge aims at providing a common…

Computer Vision and Pattern Recognition · Computer Science 2018-08-24 Abhinav Dhall , Amanjot Kaur , Roland Goecke , Tom Gedeon

The ACM Multimedia 2023 Computational Paralinguistics Challenge addresses two different problems for the first time in a research competition under well-defined conditions: In the Emotion Share Sub-Challenge, a regression on speech has to…

Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have…

Sound · Computer Science 2025-11-25 Xiangrui Xiong , Zhou Zhou , Guocai Nong , Junlin Deng , Ning Wu

The recognition of behaviors in videos usually requires a combinatorial analysis of the spatial information about objects and their dynamic action information in the temporal dimension. Specifically, behavior recognition may even rely more…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Lizong Zhang , Yiming Wang , Bei Hui , Xiujian Zhang , Sijuan Liu , Shuxin Feng

Physiological signals that provide the objective repression of human affective states are attracted increasing attention in the emotion recognition field. However, the single signal is difficult to obtain completely and accurately…

Machine Learning · Computer Science 2020-01-03 Jing Zhang , Yong Zhang , Suhua Zhan , Cheng Cheng

Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xuecheng Wu , Heli Sun , Junxiao Xue , Jiayu Nie , Xiangyan Kong , Ruofan Zhai , Liang He

This paper addresses the challenges of Online Action Recognition (OAR), a framework that involves instantaneous analysis and classification of behaviors in video streams. OAR must operate under stringent latency constraints, making it an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Wei Luo , Deyu Zhang , Ying Tang , Fan Wu , Yaoxue Zhang

Understanding emotions accurately is essential for fields like human-computer interaction. Due to the complexity of emotions and their multi-modal nature (e.g., emotions are influenced by facial expressions and audio), researchers have…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Qize Yang , Detao Bai , Yi-Xing Peng , Xihan Wei

The First Perception Test challenge was held as a half-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2023, with the goal of benchmarking state-of-the-art video models on the recently proposed…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Joseph Heyward , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

We propose a cross-modal co-attention model for continuous emotion recognition using visual-audio-linguistic information. The model consists of four blocks. The visual, audio, and linguistic blocks are used to learn the spatial-temporal…

Multimedia · Computer Science 2022-03-31 Su Zhang , Ruyi An , Yi Ding , Cuntai Guan

Human emotions entail a complex set of behavioral, physiological and cognitive changes. Current state-of-the-art models fuse the behavioral and physiological components using classic machine learning, rather than recent deep learning…

In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emotion recognition, we propose several methods for Multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Anbin QI , Zhongliang Liu , Xinyong Zhou , Jinba Xiao , Fengrun Zhang , Qi Gan , Ming Tao , Gaozheng Zhang , Lu Zhang

Video affective understanding, which aims to predict the evoked expressions by the video content, is desired for video creation and recommendation. In the recent EEV challenge, a dense affective understanding task is proposed and requires…

Computer Vision and Pattern Recognition · Computer Science 2021-06-21 Baoming Yan , Lin Wang , Ke Gao , Bo Gao , Xiao Liu , Chao Ban , Jiang Yang , Xiaobo Li

We used two multimodal models for continuous valence-arousal recognition using visual, audio, and linguistic information. The first model is the same as we used in ABAW2 and ABAW3, which employs the leader-follower attention. The second…

Multimedia · Computer Science 2023-04-18 Su Zhang , Ziyuan Zhao , Cuntai Guan

Speech emotion recognition is a challenging classification task with natural emotional speech, especially when the distribution of emotion types is imbalanced in the training and test data. In this case, it is more difficult for a model to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-31 Mingjie Chen , Hezhao Zhang , Yuanchao Li , Jiachen Luo , Wen Wu , Ziyang Ma , Peter Bell , Catherine Lai , Joshua Reiss , Lin Wang , Philip C. Woodland , Xie Chen , Huy Phan , Thomas Hain

We introduce \textbf{LongInsightBench}, the first benchmark designed to assess models' ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements, while integrating \textbf{visual,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 ZhaoYang Han , Qihan Lin , Hao Liang , Bowen Chen , Zhou Liu , Wentao Zhang