English
Related papers

Related papers: HiSpatial: Taming Hierarchical 3D Spatial Understa…

200 papers

Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xingyilang Yin , Chengzhengxu Li , Jiahao Chang , Chi-Man Pun , Xiaodong Cun

Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Jiaxin Huang , Runnan Chen , Ziwen Li , Zhengqing Gao , Xiao He , Yandong Guo , Mingming Gong , Tongliang Liu

(Renyi Qu's Master's Thesis) Recent advancements in interpretable models for vision-language tasks have achieved competitive performance; however, their interpretability often suffers due to the reliance on unstructured text outputs from…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Renyi Qu , Mark Yatskar

Vision-Language-Action (VLA) models offer promising capabilities for autonomous driving through multimodal understanding. However, their utilization in safety-critical scenarios is constrained by inherent limitations, including imprecise…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Yiru Wang , Zichong Gu , Yu Gao , Anqing Jiang , Zhigang Sun , Shuo Wang , Yuwen Heng , Hao Sun

With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Peiwen Sun , Shiqiang Lang , Dongming Wu , Yi Ding , Kaituo Feng , Huadai Liu , Zhen Ye , Rui Liu , Yun-Hui Liu , Jianan Wang , Xiangyu Yue

Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Jiyao Zhang , Mingxu Zhang , Yitong Peng , Haoxuan Liu , Chenshuo Wang , Yuxing Long , Haoyang Huang , Dongjiang Li , Nan Duan , Hui Shen , Hao Dong

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models often lack the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open challenge. Existing 3D and 4D approaches typically embed scene…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Hanyu Zhou , Gim Hee Lee

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image encoders. The…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Nahid Alam , Leema Krishna Murali , Siddhant Bharadwaj , Patrick Liu , Timothy Chung , Drishti Sharma , Akshata A , Kranthi Kiran , Wesley Tam , Bala Krishna S Vegesna

Multimodal large language models (MLLMs) have achieved significant progress in image and language tasks due to the strong reasoning capability of large language models (LLMs). Nevertheless, most MLLMs suffer from limited spatial reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Jiajie Guo , Qingpeng Zhu , Jin Zeng , Xiaolong Wu , Changyong He , Weida Wang

Text-based Visual Question Answering~(TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Hao Li , Jinfa Huang , Peng Jin , Guoli Song , Qi Wu , Jie Chen

Vision--language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-curated benchmark targeting three components of spatial understanding: depth-ordered…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Animesh Maheshwari , Divyansh Sahu , Nishit Verma

Large vision language models (LVLM) are the leading A.I approach for achieving a general visual understanding of the world. Models such as GPT, Claude, Gemini, and LLama can use images to understand and analyze complex visual scenes. 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Sagi Eppel

Large language models have shown impressive results for multi-hop mathematical reasoning when the input question is only textual. Many mathematical reasoning problems, however, contain both text and image. With the ever-increasing adoption…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Mehran Kazemi , Hamidreza Alvari , Ankit Anand , Jialin Wu , Xi Chen , Radu Soricut

Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition capabilities, most benchmarks emphasize localization…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Vahid Mirjalili , Ramin Giahi , Sriram Kollipara , Akshay Kekuda , Kehui Yao , Kai Zhao , Jianpeng Xu , Kaushiki Nag , Sinduja Subramaniam , Topojoy Biswas , Evren Korpeoglu , Kannan Achan

Most Vision-and-Language Navigation (VLN) algorithms are prone to making inaccurate decisions due to their lack of visual common sense and limited reasoning capabilities. To address this issue, we propose a Hierarchical Spatial Proximity…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Ming Xu , Zilong Xie

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Xiaoyu Zhan , Wenxuan Huang , Hao Sun , Xinyu Fu , Changfeng Ma , Shaosheng Cao , Bohan Jia , Shaohui Lin , Zhenfei Yin , Lei Bai , Wanli Ouyang , Yuanqi Li , Jie Guo , Yanwen Guo

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive and reason, showing…

Understanding geometry relies heavily on vision. In this work, we evaluate whether state-of-the-art vision language models (VLMs) can understand simple geometric concepts. We use a paradigm from cognitive science that isolates visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Eliza Kosoy , Annya Dahmani , Andrew K. Lampinen , Iulia M. Comsa , Soojin Jeong , Ishita Dasgupta , Kelsey Allen

Reasoning about fine-grained spatial relationships in warehouse-scale environments poses a significant challenge for existing vision-language models (VLMs), which often struggle to comprehend 3D layouts, object arrangements, and multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Vinh-Thuan Ly , Hoang M. Truong , Xuan-Huong Nguyen