English
Related papers

Related papers: M$^3$-ACE: Rectifying Visual Perception in Multimo…

200 papers

We present EDGE, a general-purpose, misconception-aware adaptive learning framework composed of four stages: Evaluate (ability and state estimation), Diagnose (posterior infer-ence of misconceptions), Generate (counterfactual item…

Machine Learning · Computer Science 2025-08-12 Ananda Prakash Verma

Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and integrating them effectively into the model's reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Zhuohong Chen , Zhengxian Wu , Zirui Liao , Shenao Jiang , Hangrui Xu , Yang Chen , Chaokui Su , Xiaoyu Liu , Haoqian Wang

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse vision-language tasks, yet their internal decision-making mechanisms remain insufficiently understood. Existing interpretability research has primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Jiawei Liang , Ruoyu Chen , Xianghao Jiao , Siyuan Liang , Shiming Liu , Qunli Zhang , Zheng Hu , Xiaochun Cao

Reinforcement learning (RL) has emerged as a promising approach for eliciting reasoning chains before generating final answers. However, multimodal large language models (MLLMs) generate reasoning that lacks integration of visual…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Omar Sharif , Eftekhar Hossain , Patrick Ng

Large vision-language models (VLMs) have made significant strides in 2D visual understanding tasks, sparking interest in extending these capabilities to 3D scene understanding. However, current 3D VLMs often struggle with robust reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Ting Huang , Zeyu Zhang , Hao Tang

Recent advances in multimodal large reasoning models (MLRMs) have substantially improved their ability to solve complex textual and visual tasks. However, these models tend to overthink on simple problems, producing unnecessarily lengthy…

Computation and Language · Computer Science 2025-10-10 Shuang Chen , Yue Guo , Yimeng Ye , Shijue Huang , Wenbo Hu , Haoxi Li , Manyuan Zhang , Jiayu Chen , Song Guo , Nanyun Peng

Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) demonstrate strong reasoning capabilities, yet their performance in English significantly outperforms that in low-resource languages, raising fairness concerns in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Qiming Li , Xiaocheng Feng , Yixuan Ma , Zekai Ye , Ruihan Chen , Xiachong Feng , Bing Qin

Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inputs. To bridge this gap, drawing inspiration from cognitive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jiacheng Hua , Yishu Yin , Yuhang Wu , Tai Wang , Yifei Huang , Miao Liu

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require coordinating multiple abilities, including visual…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Ziyue Wang , Yurui Dong , Fuwen Luo , Minyuan Ruan , Zhili Cheng , Chi Chen , Peng Li , Yang Liu

Recent advances in Large Language Models (LLMs) and multimodal foundation models have significantly broadened their application in robotics and collaborative systems. However, effective multi-agent interaction necessitates robust…

Video foundation models generate visually realistic and temporally coherent content, but their reliability as world simulators depends on whether they capture physical, logical, and spatial constraints. Existing metrics such as Frechet…

Computation and Language · Computer Science 2025-12-18 Zefan Cai , Haoyi Qiu , Tianyi Ma , Haozhe Zhao , Gengze Zhou , Kung-Hsiang Huang , Parisa Kordjamshidi , Minjia Zhang , Wen Xiao , Jiuxiang Gu , Nanyun Peng , Junjie Hu

Multimodal information extraction (IE) tasks have attracted increasing attention because many studies have shown that multimodal information benefits text information extraction. However, existing multimodal IE datasets mainly focus on…

Computation and Language · Computer Science 2024-12-17 Jiang Liu , Bobo Li , Xinran Yang , Na Yang , Hao Fei , Mingyao Zhang , Fei Li , Donghong Ji

Interpretability is a pressing issue for decision systems. Many post hoc methods have been proposed to explain the predictions of a single machine learning model. However, business processes and decision systems are rarely centered around a…

Machine Learning · Computer Science 2023-03-22 Gianluigi Lopardo , Damien Garreau , Frederic Precioso , Greger Ottosson

Large language models achieve strong performance on many complex reasoning tasks, yet their accuracy degrades sharply on benchmarks that require compositional reasoning, including ARC-AGI-2, GPQA, MATH, BBH, and HLE. Existing methods…

Artificial Intelligence · Computer Science 2026-02-18 Sarim Chaudhry

Medical visual question answering aims to support clinical decision-making by enabling models to answer natural language questions based on medical images. While recent advances in multi-modal learning have significantly improved…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Bo Liu , Xiangyu Zhao , Along He , Yidi Chen , Huazhu Fu , Xiao-Ming Wu

Multimodal embeddings are widely used in downstream tasks such as multimodal retrieval, enabling alignment of interleaved modalities in a shared representation space. While recent studies show that Multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Chunxu Liu , Jiyuan Yang , Ruopeng Gao , Yuhan Zhu , Feng Zhu , Rui Zhao , Limin Wang

Human communication often relies on visual cues to resolve ambiguity. While humans can intuitively integrate these cues, AI systems often find it challenging to engage in sophisticated multimodal reasoning. We introduce VAGUE, a benchmark…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Heejeong Nam , Jinwoo Ahn , Keummin Ka , Jiwan Chung , Youngjae Yu

Modern vision-language models (VLMs) deliver impressive predictive accuracy yet offer little insight into 'why' a decision is reached, frequently hallucinating facts, particularly when encountering out-of-distribution data. Neurosymbolic…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Sanchit Sinha , Guangzhi Xiong , Zhenghao He , Aidong Zhang

Even in the era of rapid advances in large models, video understanding remains a highly challenging task. Compared to texts or images, videos commonly contain more information with redundancy, requiring large models to properly allocate…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Shiwen Cao , Zhaoxing Zhang , Junming Jiao , Juyi Qiao , Guowen Song , Rong Shen , Xiangbing Meng

Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge to answer questions more accurately. However, research on evaluating RAG systems-particularly the retriever component-remains limited, as…

Information Retrieval · Computer Science 2026-04-21 Lorenz Brehme , Thomas Ströhle , Ruth Breu