English
Related papers

Related papers: MuirBench: A Comprehensive Benchmark for Robust Mu…

200 papers

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their…

Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, while spatial intelligence is fundamental for Vision-Language Models (VLMs) in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Zijian Song , Xiaoxin Lin , Qiuming Huang , Sihan Qin , Guangrun Wang , Liang Lin

With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current evaluation benchmarks predominantly focus on objective…

Computation and Language · Computer Science 2025-09-24 Haokun Li , Yazhou Zhang , Jizhi Ding , Qiuchi Li , Peng Zhang

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluations rely largely on structured-text tables or clean…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Junzhe Huang , Xiaoxiao Sun , Yan Yang , Yuxuan Hou , Ruotian Zhang , Sirui Li , Hehe Fan , Serena Yeung-Levy , Xin Yu

Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a critical bottleneck. Strikingly, MLLMs can produce correct answers even while misinterpreting crucial visual elements, masking these…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Aditya Kanade , Tanuja Ganu

Comprehending text-rich visual content is paramount for the practical application of Multimodal Large Language Models (MLLMs), since text-rich scenarios are ubiquitous in the real world, which are characterized by the presence of extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Bohao Li , Yuying Ge , Yi Chen , Yixiao Ge , Ruimao Zhang , Ying Shan

With the rapid advancement of generative models, powerful image editing methods now enable diverse and highly realistic image manipulations that far surpass traditional deepfake techniques, posing new challenges for manipulation detection.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zitong Xu , Huiyu Duan , Xiaoyu Wang , Zhaolin Cai , Kaiwei Zhang , Qiang Hu , Jing Liu , Xiongkuo Min , Guangtao Zhai

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what extent can Multimodal…

The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still face challenges in…

Computation and Language · Computer Science 2025-02-19 You Li , Heyu Huang , Chi Chen , Kaiyu Huang , Chao Huang , Zonghao Guo , Zhiyuan Liu , Jinan Xu , Yuhua Li , Ruixuan Li , Maosong Sun

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Chaoyou Fu , Peixian Chen , Yunhang Shen , Yulei Qin , Mengdan Zhang , Xu Lin , Jinrui Yang , Xiawu Zheng , Ke Li , Xing Sun , Yunsheng Wu , Rongrong Ji , Caifeng Shan , Ran He

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities, yet their proficiency in understanding and reasoning over multiple images remains largely unexplored. While existing benchmarks have initiated the evaluation of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Anurag Das , Adrian Bulat , Alberto Baldrati , Ioannis Maniadis Metaxas , Bernt Schiele , Georgios Tzimiropoulos , Brais Martinez

Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimodal math benchmarks are limited to single-visual contexts,…

Artificial Intelligence · Computer Science 2025-08-04 Peijie Wang , Zhong-Zhi Li , Fei Yin , Xin Yang , Dekang Ran , Cheng-Lin Liu

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Shuo Cao , Jiayang Li , Xiaohui Li , Yuandong Pu , Kaiwen Zhu , Yuanting Gao , Siqi Luo , Yi Xin , Qi Qin , Yu Zhou , Xiangyu Chen , Wenlong Zhang , Bin Fu , Yu Qiao , Yihao Liu

Multimodal Large Language Models (MLLMs) have shown impressive abilities in understanding and reasoning over conventional images. However, their perception of 360{\deg} images remains largely underexplored. Unlike conventional images,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Huyen T. T. Tran , Van-Quang Nguyen , Farros Alferro , Kang-Jun Liu , Takayuki Okatani

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Junjie Yang , Yuhao Yan , Gang Wu , Yuxuan Wang , Ruoyu Liang , Xinjie Jiang , Xiang Wan , Fenglei Fan , Yongquan Zhang , Feiwei Qin , Changmiao Wang

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuansen Liu , Haiming Tang , Jinlong Peng , Jiangning Zhang , Xiaozhong Ji , Qingdong He , Wenbin Wu , Donghao Luo , Zhenye Gan , Junwei Zhu , Yunhang Shen , Chaoyou Fu , Chengjie Wang , Xiaobin Hu , Shuicheng Yan

The security concerns surrounding Large Language Models (LLMs) have been extensively explored, yet the safety of Multimodal Large Language Models (MLLMs) remains understudied. In this paper, we observe that Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Xin Liu , Yichen Zhu , Jindong Gu , Yunshi Lan , Chao Yang , Yu Qiao

With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Micro-Action understanding, a vital role in human emotion analysis, remains unexplored due to the absence of specialized benchmarks. To tackle this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Kun Li , Jihao Gu , Fei Wang , Zhiliang Wu , Hehe Fan , Dan Guo

Large Vision-Language Models (LVLMs) and Multimodal Large Language Models (MLLMs) have demonstrated outstanding performance in various general multimodal applications and have shown increasing promise in specialized domains. However, their…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Chenwei Lin , Hanjia Lyu , Xian Xu , Jiebo Luo

Though Multi-modal Large Language Models (MLLMs) have recently achieved significant progress, they often struggle to understand diverse and complicated inter-object relations. Specifically, the lack of large-scale and high-quality relation…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jiahao Nie , Gongjie Zhang , Wenbin An , Yun Xing , Yap-Peng Tan , Alex C. Kot , Shijian Lu