中文
相关论文

相关论文: OMG - Emotion Challenge Solution

200 篇论文

Much of the appeal of music lies in its power to convey emotions/moods and to evoke them in listeners. In consequence, the past decade witnessed a growing interest in modeling emotions from musical signals in the music information retrieval…

信息检索 · 计算机科学 2015-02-19 Ju-Chiang Wang , Yi-Hsuan Yang , Hsin-Min Wang

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

Dynamic emotion recognition in the wild remains challenging due to the transient nature of emotional expressions and temporal misalignment of multi-modal cues. Traditional approaches predict valence and arousal and often overlook the…

Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Zhe Kong , Yong Zhang , Tianyu Yang , Tao Wang , Kaihao Zhang , Bizhu Wu , Guanying Chen , Wei Liu , Wenhan Luo

This paper illustrates our submission method to the fourth Affective Behavior Analysis in-the-Wild (ABAW) Competition. The method is used for the Multi-Task Learning Challenge. Instead of using only face information, we employ full…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Irfan Haider , Minh-Trieu Tran , Soo-Hyung Kim , Hyung-Jeong Yang , Guee-Sang Lee

This paper introduces the YouTube-8M Video Understanding Challenge hosted as a Kaggle competition and also describes my approach to experimenting with various models. For each of my experiments, I provide the score result as well as…

机器学习 · 统计学 2017-06-27 Edward Chen

Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, often treat emotion recognition as static fusion over…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Bo Zhao , Fanghua Ye , Yixin Ji , Sicheng Zhao , Xiaojiang Peng , Zitong YU

We propose a novel framework for open-ended video question answering that enhances reasoning depth and robustness in complex real-world scenarios, as benchmarked on the CVRR-ES dataset. Existing Video-Large Multimodal Models (Video-LMMs)…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Jun Xie , Zhaoran Zhao , Xiongjun Guan , Yingjian Zhu , Hongzhu Yi , Xinming Wang , Feng Chen , Zhepeng Wang

The state of the art in video understanding suffers from two problems: (1) The major part of reasoning is performed locally in the video, therefore, it misses important relationships within actions that span several seconds. (2) While there…

计算机视觉与模式识别 · 计算机科学 2018-05-08 Mohammadreza Zolfaghari , Kamaljeet Singh , Thomas Brox

Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Jun Yu , Lingsi Zhu , Yanjun Chi , Yunxiang Zhang , Yang Zheng , Yongqi Wang , Xilong Lu

Emotion Representation Mapping (ERM) has the goal to convert existing emotion ratings from one representation format into another one, e.g., mapping Valence-Arousal-Dominance annotations for words or sentences into Ekman's Basic Emotions…

计算与语言 · 计算机科学 2018-06-26 Sven Buechel , Udo Hahn

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current…

From a computational viewpoint, emotions continue to be intriguingly hard to understand. In research, direct, real-time inspection in realistic settings is not possible. Discrete, indirect, post-hoc recordings are therefore the norm. As a…

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, including Moment Queries, Natural Language Queries, Future Hand…

Emotion recognition and classification is a very active area of research. In this paper, we present a first approach to emotion classification using persistent entropy and support vector machines. A topology-based model is applied to obtain…

声音 · 计算机科学 2019-03-22 R. Gonzalez-Diaz , E. Paluzo-Hidalgo , J. F. Quesada

Emotion recognition from speech is a challenging task that requires capturing both linguistic and paralinguistic cues, with critical applications in human-computer interaction and mental health monitoring. Recent works have highlighted the…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Hugo Thimonier , Antony Perzo , Renaud Seguier

The conditional moment problem is a powerful formulation for describing structural causal parameters in terms of observables, a prominent example being instrumental variable regression. A standard approach reduces the problem to a finite…

机器学习 · 计算机科学 2023-03-24 Andrew Bennett , Nathan Kallus

This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Quality Assessment (KVQ), and (ii) Diffusion-based Image…

图像与视频处理 · 电气工程与系统科学 2025-04-18 Xin Li , Kun Yuan , Bingchen Li , Fengbin Guan , Yizhen Shao , Zihao Yu , Xijun Wang , Yiting Lu , Wei Luo , Suhang Yao , Ming Sun , Chao Zhou , Zhibo Chen , Radu Timofte , Yabin Zhang , Ao-Xiang Zhang , Tianwu Zhi , Jianzhao Liu , Yang Li , Jingwen Xu , Yiting Liao , Yushen Zuo , Mingyang Wu , Renjie Li , Shengyun Zhong , Zhengzhong Tu , Yufan Liu , Xiangguang Chen , Zuowei Cao , Minhao Tang , Shan Liu , Kexin Zhang , Jingfen Xie , Yan Wang , Kai Chen , Shijie Zhao , Yunchen Zhang , Xiangkai Xu , Hong Gao , Ji Shi , Yiming Bao , Xiugang Dong , Xiangsheng Zhou , Yaofeng Tu , Ying Liang , Yiwen Wang , Xinning Chai , Yuxuan Zhang , Zhengxue Cheng , Yingsheng Qin , Yucai Yang , Rong Xie , Li Song , Wei Sun , Kang Fu , Linhan Cao , Dandan Zhu , Kaiwei Zhang , Yucheng Zhu , Zicheng Zhang , Menghan Hu , Xiongkuo Min , Guangtao Zhai , Zhi Jin , Jiawei Wu , Wei Wang , Wenjian Zhang , Yuhai Lan , Gaoxiong Yi , Hengyuan Na , Wang Luo , Di Wu , MingYin Bai , Jiawang Du , Zilong Lu , Zhenyu Jiang , Hui Zeng , Ziguan Cui , Zongliang Gan , Guijin Tang , Xinglin Xie , Kehuan Song , Xiaoqiang Lu , Licheng Jiao , Fang Liu , Xu Liu , Puhua Chen , Ha Thu Nguyen , Katrien De Moor , Seyed Ali Amirshahi , Mohamed-Chaker Larabi , Qi Tang , Linfeng He , Zhiyong Gao , Zixuan Gao , Guohua Zhang , Zhiye Huang , Yi Deng , Qingmiao Jiang , Lu Chen , Yi Yang , Xi Liao , Nourine Mohammed Nadir , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Meiqin Liu , Chao Yao , Yao Zhao

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision or video models and a move towards long-video understanding.…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Darshana Saravanan , Varun Gupta , Darshan Singh , Zeeshan Khan , Vineet Gandhi , Makarand Tapaswi

Online continual learning aims to get closer to a live learning experience by learning directly on a stream of data with temporally shifting distribution and by storing a minimum amount of data from that stream. In this empirical…