English
Related papers

Related papers: Thinking in Structures: Evaluating Spatial Intelli…

200 papers

Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"? We curate three new corpora to quantify model comprehension of such basic spatial relations. These tests isolate spatial reasoning more…

Computation and Language · Computer Science 2023-10-31 Amita Kamath , Jack Hessel , Kai-Wei Chang

Large language models have shown impressive results for multi-hop mathematical reasoning when the input question is only textual. Many mathematical reasoning problems, however, contain both text and image. With the ever-increasing adoption…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Mehran Kazemi , Hamidreza Alvari , Ankit Anand , Jialin Wu , Xi Chen , Radu Soricut

Large vision-and-language models (VLMs) trained to match images with text on large-scale datasets of image-text pairs have shown impressive generalization ability on several vision and language tasks. Several recent works, however, showed…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Navid Rajabi , Jana Kosecka

Humans build shared spatial understanding by communicating partial, viewpoint-dependent observations. We ask whether Multimodal Large Language Models (MLLMs) can do the same, aligning distinct egocentric views through dialogue to form a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Ankur Sikarwar , Debangan Mishra , Sudarshan Nikhil , Ponnurangam Kumaraguru , Aishwarya Agrawal

Solid geometry problem solving demands spatial mathematical reasoning that integrates spatial intelligence and symbolic reasoning. However, most existing multimodal mathematical reasoning benchmarks focus primarily on 2D plane geometry,…

Artificial Intelligence · Computer Science 2025-11-12 Changti Wu , Shijie Lian , Zihao Liu , Lei Zhang , Laurence Tianruo Yang , Kai Chen

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhibit significant…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Kexin Tian , Jingrui Mao , Yunlong Zhang , Jiwan Jiang , Yang Zhou , Zhengzhong Tu

We introduce STSBench, a scenario-based framework to benchmark the holistic understanding of vision-language models (VLMs) for autonomous driving. The framework automatically mines pre-defined traffic scenarios from any dataset using…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Christian Fruhwirth-Reisinger , Dušan Malić , Wei Lin , David Schinagl , Samuel Schulter , Horst Possegger

Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. Due to extensive training on vast amounts of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Pranshu Pandya , Vatsal Gupta , Agney S Talwarr , Tushar Kataria , Dan Roth , Vivek Gupta

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning…

Spatial understanding is fundamental for embodied agents, yet most spatial VLMs and benchmarks remain offline-evaluating post-hoc QA over pre-recorded inputs and overlooking two crucial deployment-critical requirements: long-horizon…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Yuxi Wei , Wei Huang , Qirui Chen , Lu Hou , Xiaojuan Qi

With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Peiwen Sun , Shiqiang Lang , Dongming Wu , Yi Ding , Kaituo Feng , Huadai Liu , Zhen Ye , Rui Liu , Yun-Hui Liu , Jianan Wang , Xiangyu Yue

Text-based Visual Question Answering~(TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Hao Li , Jinfa Huang , Peng Jin , Guoli Song , Qi Wu , Jie Chen

Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Most existing works improve spatial reasoning by introducing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhenghao Chen , Huiqun Wang , Di Huang

When faced with complex spatial problems, humans naturally sketch layouts to organize their thinking, and the act of drawing further sharpens their understanding. In this work, we ask whether a similar principle holds for Large Language…

Artificial Intelligence · Computer Science 2026-04-17 Shiyuan Huang , Li Liu , Jincheng He , Leilani H. Gilpin

Spatial visualization is the mental ability to imagine, transform, and manipulate the spatial characteristics of objects and actions. This intelligence is a part of human cognition where actions and perception are connected on a mental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Nilay Yilmaz , Maitreya Patel , Naga Sai Abhiram Kusumba , Yixuan He , Yezhou Yang

Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. We present…

Machine Learning · Computer Science 2025-06-03 Xinwu Ye , Chengfan Li , Siming Chen , Wei Wei , Xiangru Tang

Despite impressive high-level video comprehension, multimodal language models struggle with spatial reasoning across time and space. While current spatial training approaches rely on real-world video data, obtaining diverse footage with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Ellis Brown , Arijit Ray , Ranjay Krishna , Ross Girshick , Rob Fergus , Saining Xie

Large language models (LLMs) can perform reasoning computations both internally within their latent space and externally by generating explicit token sequences like chains of thought. Significant progress in enhancing reasoning abilities…

Computation and Language · Computer Science 2025-04-16 Thilo Hagendorff , Sarah Fabi

As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behaviors. Yet, existing multimodal benchmarks either…

Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference across various subjects. Existing benchmarks for multimodal large language models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Longteng Guo , Xuanxu Lin , Dongze Hao , Tongtian Yue , Pengkang Huo , Jiatong Ma , Yuchen Liu , Jing Liu