English
Related papers

Related papers: Learning What to Attend First: Modality-Importance…

200 papers

Recent advances in large language models (LLMs) have opened new avenues for multimodal reasoning. Yet, most existing methods still rely on pretrained vision-language models (VLMs) to encode image-text pairs in isolation, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yuanfu Sun , Kang Li , Pengkang Guo , Jiajin Liu , Qiaoyu Tan

Multimodal emotion recognition in conversation (MERC) and multimodal emotion-cause pair extraction (MECPE) have recently garnered significant attention. Emotions are the expression of affect or feelings; responses to specific events, or…

Computation and Language · Computer Science 2024-10-10 Guimin Hu , Zhihong Zhu , Daniel Hershcovich , Lijie Hu , Hasti Seifi , Jiayuan Xie

We propose a novel method, Modality-based Redundancy Reduction Fusion (MRRF), for understanding and modulating the relative contribution of each modality in multimodal inference tasks. This is achieved by obtaining an $(M+1)$-way tensor to…

Machine Learning · Computer Science 2023-04-18 Elham J. Barezi , Peyman Momeni , Pascale Fung

Multi-modal models have shown a promising capability to effectively integrate information from various sources, yet meanwhile, they are found vulnerable to pervasive perturbations, such as uni-modal attacks and missing conditions. To…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Zequn Yang , Yake Wei , Ce Liang , Di Hu

Recent advances in motion-aware large language models have shown remarkable promise for unifying motion understanding and generation tasks. However, these models typically treat understanding and generation separately, limiting the mutual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Yuan-Ming Li , Qize Yang , Nan Lei , Shenghao Fu , Ling-An Zeng , Jian-Fang Hu , Xihan Wei , Wei-Shi Zheng

Dynamic Facial Expression Recognition (DFER) aims to identify human emotions from temporally evolving facial movements and plays a critical role in affective computing. While recent vision-language approaches have introduced semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Yu Liu , Leyuan Qu , Hanlei Shi , Di Gao , Yuhua Zheng , Taihao Li

Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse tasks. However, their deployment in long context scenarios remains hindered by computational inefficiency and information redundancy. Context compression…

Computation and Language · Computer Science 2026-03-09 Jiwei Tang , Shilei Liu , Zhicheng Zhang , Yujin Yuan , Libin Zheng , Wenbo Su , Bo Zheng

Biological function emerges from coupled constraints across sequence, structure, regulation, evolution, and cellular context, yet most foundation models in biology are trained within one modality or for a fixed forward task. We present…

Despite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous presence of all involved multiple modalities. Motivated by…

Machine Learning · Computer Science 2019-07-25 Jing Han , Zixing Zhang , Zhao Ren , Björn Schuller

While recent large language models (LLMs) improve on various question answering (QA) datasets, it remains difficult for a single model to generalize across question types that require distinct reasoning abilities. We provide empirical…

Computation and Language · Computer Science 2023-10-23 Chenglei Si , Weijia Shi , Chen Zhao , Luke Zettlemoyer , Jordan Boyd-Graber

Multimodal Object-Entity Relation Extraction (MORE) is a challenging task in information extraction research. It aims to identify relations between visual objects and textual entities, requiring complex multimodal understanding and…

Multimedia · Computer Science 2026-03-11 Xiang Yuan , Xu Chu , Xinrong Chen , Haochen Li , Zonghong Dai , Hongcheng Fan , Xiaoyue Yuan , Weiping Li , Tong Mo

Medical decision systems increasingly rely on data from multiple sources to ensure reliable and unbiased diagnosis. However, existing multimodal learning models fail to achieve this goal because they often ignore two critical challenges.…

Machine Learning · Computer Science 2025-10-10 Md Zubair , Hao Zheng , Nussdorf Jonathan , Grayson W. Armstrong , Lucy Q. Shen , Gabriela Wilson , Yu Tian , Xingquan Zhu , Min Shi

In this work, we present the first application of Reinforcement Learning with Verifiable Reward (RLVR) to an Omni-multimodal large language model in the context of emotion recognition, a task where both visual and audio modalities play…

Machine Learning · Computer Science 2025-03-11 Jiaxing Zhao , Xihan Wei , Liefeng Bo

Recent advances in synergizing large reasoning models (LRMs) with retrieval-augmented generation (RAG) have shown promising results, yet two critical challenges remain: (1) reasoning models typically operate from a single, unchallenged…

Artificial Intelligence · Computer Science 2026-01-12 Can Xu , Lingyong Yan , Jiayi Wu , Haosen Wang , Shuaiqiang Wang , Yuchen Li , Jizhou Huang , Dawei Yin , Xiang Li

Although large language models (LLMs) have demonstrated remarkable reasoning capabilities, they still face challenges in knowledge-intensive multi-hop reasoning. Recent work explores iterative retrieval to address complex problems. However,…

Computation and Language · Computer Science 2025-05-27 Zheng Chu , Huiming Fan , Jingchang Chen , Qianyu Wang , Mingda Yang , Jiafeng Liang , Zhongjie Wang , Hao Li , Guo Tang , Ming Liu , Bing Qin

Large Language Models (LLMs) have shown impressive capabilities in natural language processing but still struggle to perform well on knowledge-intensive tasks that require deep reasoning and the integration of external knowledge. Although…

Computation and Language · Computer Science 2025-08-26 Bo Zhao , Yinghao Zhang , Ziqi Xu , Yongli Ren , Xiuzhen Zhang , Renqiang Luo , Zaiwen Feng , Feng Xia

Multimodal Retrieval-Augmented Generation (RAG) systems have become essential in knowledge-intensive and open-domain tasks. As retrieval complexity increases, ensuring the robustness of these systems is critical. However, current RAG models…

Computation and Language · Computer Science 2025-06-16 Jiayu Yao , Shenghua Liu , Yiwei Wang , Lingrui Mei , Baolong Bi , Yuyao Ge , Zhecheng Li , Xueqi Cheng

Self-consistency methods are the core technique for improving the reasoning reliability of multimodal large language models (MLLMs). By generating multiple reasoning results through repeated sampling and selecting the best answer via…

Computation and Language · Computer Science 2026-02-05 Xinglong Yang , Zhilin Peng , Zhanzhan Liu , Haochen Shi , Sheng-Jun Huang

Multimodal sentiment analysis is a core research area that studies speaker sentiment expressed from the language, visual, and acoustic modalities. The central challenge in multimodal learning involves inferring joint representations that…

Machine Learning · Computer Science 2020-03-02 Hai Pham , Paul Pu Liang , Thomas Manzini , Louis-Philippe Morency , Barnabas Poczos

Multimodal emotion recognition study is hindered by the lack of labelled corpora in terms of scale and diversity, due to the high annotation cost and label ambiguity. In this paper, we propose a pre-training model \textbf{MEmoBERT} for…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Jinming Zhao , Ruichen Li , Qin Jin , Xinchao Wang , Haizhou Li
‹ Prev 1 8 9 10 Next ›