English
Related papers

Related papers: OMG - Emotion Challenge Solution

200 papers

Much of the appeal of music lies in its power to convey emotions/moods and to evoke them in listeners. In consequence, the past decade witnessed a growing interest in modeling emotions from musical signals in the music information retrieval…

Information Retrieval · Computer Science 2015-02-19 Ju-Chiang Wang , Yi-Hsuan Yang , Hsin-Min Wang

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

Dynamic emotion recognition in the wild remains challenging due to the transient nature of emotional expressions and temporal misalignment of multi-modal cues. Traditional approaches predict valence and arousal and often overlook the…

Machine Learning · Computer Science 2025-05-05 Vrushank Ahire , Kunal Shah , Mudasir Nazir Khan , Nikhil Pakhale , Lownish Rai Sookha , M. A. Ganaie , Abhinav Dhall

Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Zhe Kong , Yong Zhang , Tianyu Yang , Tao Wang , Kaihao Zhang , Bizhu Wu , Guanying Chen , Wei Liu , Wenhan Luo

This paper illustrates our submission method to the fourth Affective Behavior Analysis in-the-Wild (ABAW) Competition. The method is used for the Multi-Task Learning Challenge. Instead of using only face information, we employ full…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Irfan Haider , Minh-Trieu Tran , Soo-Hyung Kim , Hyung-Jeong Yang , Guee-Sang Lee

This paper introduces the YouTube-8M Video Understanding Challenge hosted as a Kaggle competition and also describes my approach to experimenting with various models. For each of my experiments, I provide the score result as well as…

Machine Learning · Statistics 2017-06-27 Edward Chen

Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, often treat emotion recognition as static fusion over…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Bo Zhao , Fanghua Ye , Yixin Ji , Sicheng Zhao , Xiaojiang Peng , Zitong YU

We propose a novel framework for open-ended video question answering that enhances reasoning depth and robustness in complex real-world scenarios, as benchmarked on the CVRR-ES dataset. Existing Video-Large Multimodal Models (Video-LMMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Jun Xie , Zhaoran Zhao , Xiongjun Guan , Yingjian Zhu , Hongzhu Yi , Xinming Wang , Feng Chen , Zhepeng Wang

The state of the art in video understanding suffers from two problems: (1) The major part of reasoning is performed locally in the video, therefore, it misses important relationships within actions that span several seconds. (2) While there…

Computer Vision and Pattern Recognition · Computer Science 2018-05-08 Mohammadreza Zolfaghari , Kamaljeet Singh , Thomas Brox

Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Jun Yu , Lingsi Zhu , Yanjun Chi , Yunxiang Zhang , Yang Zheng , Yongqi Wang , Xilong Lu

Emotion Representation Mapping (ERM) has the goal to convert existing emotion ratings from one representation format into another one, e.g., mapping Valence-Arousal-Dominance annotations for words or sentences into Ekman's Basic Emotions…

Computation and Language · Computer Science 2018-06-26 Sven Buechel , Udo Hahn

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current…

From a computational viewpoint, emotions continue to be intriguingly hard to understand. In research, direct, real-time inspection in realistic settings is not possible. Discrete, indirect, post-hoc recordings are therefore the norm. As a…

Human-Computer Interaction · Computer Science 2018-12-10 Karan Sharma , Claudio Castellini , Egon L. van den Broek , Alin Albu-Schaeffer , Friedhelm Schwenker

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, including Moment Queries, Natural Language Queries, Future Hand…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Guo Chen , Sen Xing , Zhe Chen , Yi Wang , Kunchang Li , Yizhuo Li , Yi Liu , Jiahao Wang , Yin-Dong Zheng , Bingkun Huang , Zhiyu Zhao , Junting Pan , Yifei Huang , Zun Wang , Jiashuo Yu , Yinan He , Hongjie Zhang , Tong Lu , Yali Wang , Limin Wang , Yu Qiao

Emotion recognition and classification is a very active area of research. In this paper, we present a first approach to emotion classification using persistent entropy and support vector machines. A topology-based model is applied to obtain…

Sound · Computer Science 2019-03-22 R. Gonzalez-Diaz , E. Paluzo-Hidalgo , J. F. Quesada

Emotion recognition from speech is a challenging task that requires capturing both linguistic and paralinguistic cues, with critical applications in human-computer interaction and mental health monitoring. Recent works have highlighted the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Hugo Thimonier , Antony Perzo , Renaud Seguier

The conditional moment problem is a powerful formulation for describing structural causal parameters in terms of observables, a prominent example being instrumental variable regression. A standard approach reduces the problem to a finite…

Machine Learning · Computer Science 2023-03-24 Andrew Bennett , Nathan Kallus

This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Quality Assessment (KVQ), and (ii) Diffusion-based Image…

Image and Video Processing · Electrical Eng. & Systems 2025-04-18 Xin Li , Kun Yuan , Bingchen Li , Fengbin Guan , Yizhen Shao , Zihao Yu , Xijun Wang , Yiting Lu , Wei Luo , Suhang Yao , Ming Sun , Chao Zhou , Zhibo Chen , Radu Timofte , Yabin Zhang , Ao-Xiang Zhang , Tianwu Zhi , Jianzhao Liu , Yang Li , Jingwen Xu , Yiting Liao , Yushen Zuo , Mingyang Wu , Renjie Li , Shengyun Zhong , Zhengzhong Tu , Yufan Liu , Xiangguang Chen , Zuowei Cao , Minhao Tang , Shan Liu , Kexin Zhang , Jingfen Xie , Yan Wang , Kai Chen , Shijie Zhao , Yunchen Zhang , Xiangkai Xu , Hong Gao , Ji Shi , Yiming Bao , Xiugang Dong , Xiangsheng Zhou , Yaofeng Tu , Ying Liang , Yiwen Wang , Xinning Chai , Yuxuan Zhang , Zhengxue Cheng , Yingsheng Qin , Yucai Yang , Rong Xie , Li Song , Wei Sun , Kang Fu , Linhan Cao , Dandan Zhu , Kaiwei Zhang , Yucheng Zhu , Zicheng Zhang , Menghan Hu , Xiongkuo Min , Guangtao Zhai , Zhi Jin , Jiawei Wu , Wei Wang , Wenjian Zhang , Yuhai Lan , Gaoxiong Yi , Hengyuan Na , Wang Luo , Di Wu , MingYin Bai , Jiawang Du , Zilong Lu , Zhenyu Jiang , Hui Zeng , Ziguan Cui , Zongliang Gan , Guijin Tang , Xinglin Xie , Kehuan Song , Xiaoqiang Lu , Licheng Jiao , Fang Liu , Xu Liu , Puhua Chen , Ha Thu Nguyen , Katrien De Moor , Seyed Ali Amirshahi , Mohamed-Chaker Larabi , Qi Tang , Linfeng He , Zhiyong Gao , Zixuan Gao , Guohua Zhang , Zhiye Huang , Yi Deng , Qingmiao Jiang , Lu Chen , Yi Yang , Xi Liao , Nourine Mohammed Nadir , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Meiqin Liu , Chao Yao , Yao Zhao

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision or video models and a move towards long-video understanding.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Darshana Saravanan , Varun Gupta , Darshan Singh , Zeeshan Khan , Vineet Gandhi , Makarand Tapaswi

Online continual learning aims to get closer to a live learning experience by learning directly on a stream of data with temporally shifting distribution and by storing a minimum amount of data from that stream. In this empirical…