English
Related papers

Related papers: SpatiaLab: Can Vision-Language Models Perform Spat…

200 papers

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks…

Artificial Intelligence · Computer Science 2026-05-08 Peiran Xu , Sudong Wang , Yao Zhu , Jianing Li , Gege Qi , Yunjian Zhang

Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning directly grounded in visual evidence. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Longteng Guo , Yifan Wang , Pengkang Huo , Tailai Chen , Yuze Wu , Jing Liu , Xinxin Zhu

Spatio-temporal reasoning is essential in understanding real-world environments in various fields, eg, autonomous driving and sports analytics. Recent advances have improved the spatial reasoning ability of Vision-Language Models (VLMs) by…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Dohwan Ko , Sihyeon Kim , Yumin Suh , Vijay Kumar B. G , Minseo Yoon , Manmohan Chandraker , Hyunwoo J. Kim

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Minkyu Kim , Sangheon Lee , Dongmin Park

Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations,…

For human cognitive process, spatial reasoning and perception are closely entangled, yet the nature of this interplay remains underexplored in the evaluation of multimodal large language models (MLLMs). While recent MLLM advancements show…

Computation and Language · Computer Science 2025-08-28 Chengzu Li , Wenshan Wu , Huanyu Zhang , Qingtao Li , Zeyu Gao , Yan Xia , José Hernández-Orallo , Ivan Vulić , Furu Wei

Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single-step perception-to-judgment…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Rui Zhu , Xin Shen , Shuchen Wu , Chenxi Miao , Xin Yu , Yang Li , Weikang Li , Deguo Xia , Jizhou Huang

Benchmarking spatial reasoning in multimodal large language models (MLLMs) has attracted growing interest in computer vision due to its importance for embodied AI and other agentic systems that require precise interaction with the physical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Zelin Xu , Yupu Zhang , Saugat Adhikari , Saiful Islam , Tingsong Xiao , Zibo Liu , Shigang Chen , Da Yan , Zhe Jiang

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

In this article, we investigate vision-language models (VLM) as reasoners. The ability to form abstractions underlies mathematical reasoning, problem-solving, and other Math AI tasks. Several formalisms have been given to these underlying…

Artificial Intelligence · Computer Science 2024-07-08 Denisa Roberts , Lucas Roberts

Spatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language…

Computation and Language · Computer Science 2025-04-18 Zheyuan Zhang , Fengyuan Hu , Jayjun Lee , Freda Shi , Parisa Kordjamshidi , Joyce Chai , Ziqiao Ma

Spatial reasoning plays a vital role in both human cognition and machine intelligence, prompting new research into language models' (LMs) capabilities in this regard. However, existing benchmarks reveal shortcomings in evaluating…

Computation and Language · Computer Science 2024-05-27 Fangjun Li , David C. Hogg , Anthony G. Cohn

Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but they continue to struggle with spatial understanding. Existing spatial MLLMs often rely on explicit 3D inputs or architecture-specific…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Hunar Batra , Haoqin Tu , Hardy Chen , Yuanze Lin , Cihang Xie , Ronald Clark

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high level relationships ('left…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Vineet Bhat , Sungsu Kim , Valts Blukis , Greg Heinrich , Prashanth Krishnamurthy , Ramesh Karri , Stan Birchfield , Farshad Khorrami , Jonathan Tremblay

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning,…

Machine Learning · Computer Science 2026-01-27 Ashutosh Bajpai , Akshat Bhandari , Akshay Nambi , Tanmoy Chakraborty

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Wufei Ma , Luoxin Ye , Celso M de Melo , Jieneng Chen , Alan Yuille

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Huizhi Liang , Yichao Shen , Yu Deng , Sicheng Xu , Zhiyuan Feng , Tong Zhang , Yaobo Liang , Jiaolong Yang

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become increasingly pivotal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Jiacheng Ruan , Wenzhen Yuan , Xian Gao , Ye Guo , Daoxin Zhang , Zhe Xu , Yao Hu , Ting Liu , Yuzhuo Fu

We present MetaSpatial, the first reinforcement learning (RL)-based framework designed to enhance 3D spatial reasoning in vision-language models (VLMs), enabling real-time 3D scene generation without the need for hard-coded optimizations.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Zhenyu Pan , Han Liu

Visual spatial intelligence is critical for medical image interpretation, yet remains largely unexplored in Multimodal Large Language Models (MLLMs) for 3D imaging. This gap persists due to a systemic lack of datasets featuring structured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Quoc-Huy Trinh , Xi Ding , Yang Liu , Zhenyue Qin , Xingjian Li , Gorkem Durak , Halil Ertugrul Aktas , Elif Keles , Ulas Bagci , Min Xu
‹ Prev 1 3 4 5 6 7 10 Next ›