English
Related papers

Related papers: World2Mind: Cognition Toolkit for Allocentric Spat…

200 papers

Recent test-time reasoning methods improve performance by generating more candidate chains or searching over larger reasoning trees, but they typically lack explicit control over when to expand, what to prune, how to repair, and when to…

Artificial Intelligence · Computer Science 2026-03-31 Siyuan Ma , Bo Gao , Zikai Xiao , Hailong Wang , Xinlei Yu , Rui Qian , Jiayu Qian , Luqi Gong , Yang Liu

Spatial intelligence spans a rich suite of abilities, including visualising and transforming shapes, mentally rotating objects, judging relational positions and containment, and estimating numerosity. However, it still remains a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Shijie Lian , Changti Wu , Laurence Tianruo Yang , Hang Yuan , Bin Yu , Lei Zhang , Kai Chen

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the Grounded-Spatial Reasoner (GS-Reasoner) to explore the effective spatial representations that bridge the gap between them. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yiming Chen , Zekun Qi , Wenyao Zhang , Xin Jin , Li Zhang , Peidong Liu

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Chan Hee Song , Valts Blukis , Jonathan Tremblay , Stephen Tyree , Yu Su , Stan Birchfield

Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains significantly limited. Current methods often rely on passive…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Meng Cao , Xingyu Li , Xue Liu , Ian Reid , Xiaodan Liang

While contemporary Vision-Language Models (VLMs) excel at 2D visual understanding, they remain constrained by a passive, 2D-centric paradigm that severely limits genuine 3D spatial reasoning. To bridge this gap, we introduce Think3D, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zaibin Zhang , Yuhan Wu , Lianjie Jia , Yifan Wang , Zhongbo Zhang , Yijiang Li , Binghao Ran , Fuxi Zhang , Zhuohan Sun , Zhenfei Yin , Lijun Wang , Huchuan Lu

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

Spatial reasoning in large language models (LLMs) has gained increasing attention due to applications in navigation and planning. Despite strong general language capabilities, LLMs still struggle with spatial transformations and multi-step…

Artificial Intelligence · Computer Science 2026-01-01 Amir Tahmasbi , Sadegh Majidi , Kazem Taram , Aniket Bera

While diffusion models have shown exceptional capabilities in aesthetic image synthesis, they often struggle with complex spatial understanding and reasoning. Existing approaches resort to Multimodal Large Language Models (MLLMs) to enhance…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Wei Chen , Yancheng Long , Mingqiao Liu , Haojie Ding , Yankai Yang , Hongyang Wei , Yi-Fan Zhang , Bin Wen , Fan Yang , Tingting Gao , Han Li , Long Chen

For human cognitive process, spatial reasoning and perception are closely entangled, yet the nature of this interplay remains underexplored in the evaluation of multimodal large language models (MLLMs). While recent MLLM advancements show…

Computation and Language · Computer Science 2025-08-28 Chengzu Li , Wenshan Wu , Huanyu Zhang , Qingtao Li , Zeyu Gao , Yan Xia , José Hernández-Orallo , Ivan Vulić , Furu Wei

While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this gap, we introduce geometrically referenced 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jiangye Yuan , Gowri Kumar , Baoyuan Wang

Multimodal Reasoning Models (MRMs) leveraging Chain-of-Thought (CoT) based thinking have revolutionized mathematical and logical problem-solving. However, we show that this paradigm struggles with generalized spatial intelligence. We…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Sai Srinivas Kancheti , Aditya Sanjiv Kanade , Vineeth N. Balasubramanian , Tanuja Ganu

Recent advances in Omni models have enabled unified multimodal perception and generation. However, most existing systems still exhibit rigid reasoning behaviors, either overthinking simple problems or failing to reason when necessary. To…

Artificial Intelligence · Computer Science 2025-12-05 Dongchao Yang , Songxiang Liu , Disong Wang , Yuanyuan Wang , Guanglu Wan , Helen Meng

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 An-Chieh Cheng , Hongxu Yin , Yang Fu , Qiushan Guo , Ruihan Yang , Jan Kautz , Xiaolong Wang , Sifei Liu

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional…

Applying AI foundation models directly to geospatial datasets remains challenging due to their limited ability to represent and reason with geographical entities, specifically vector-based geometries and natural language descriptions of…

Computation and Language · Computer Science 2025-05-26 Yuhan Ji , Song Gao , Ying Nie , Ivan Majić , Krzysztof Janowicz

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Zuntao Liu , Yi Du , Taimeng Fu , Shaoshu Su , Cherie Ho , Chen Wang

Precise spatial understanding from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs), as their visual representations are predominantly semantic and lack explicit geometric grounding. While…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Chanyoung Gwak , Yoonwoo Jeong , Byungwoo Jeon , Hyunseok Lee , Jinwoo Shin , Minsu Cho

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Wufei Ma , Haoyu Chen , Guofeng Zhang , Yu-Cheng Chou , Jieneng Chen , Celso M de Melo , Alan Yuille

Geolocation, the task of identifying the geographic location of an image, requires abundant world knowledge and complex reasoning abilities. Though advanced large multimodal models (LMMs) have shown superior aforementioned capabilities,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Yushuo Zheng , Huiyu Duan , Zicheng Zhang , Xiaohong Liu , Xiongkuo Min