English
Related papers

Related papers: From Where Things Are to What They Are For: Benchm…

200 papers

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are…

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

Reasoning is a fundamental cognitive process that enables logical inference, problem-solving, and decision-making. With the rapid advancement of large language models (LLMs), reasoning has emerged as a key capability that distinguishes…

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Mohsen Gholami , Ahmad Rezaei , Zhou Weimin , Sitong Mao , Shunbo Zhou , Yong Zhang , Mohammad Akbari

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiahui Zhang , Yurui Chen , Yanpeng Zhou , Yueming Xu , Ze Huang , Jilin Mei , Junhui Chen , Yu-Jie Yuan , Xinyue Cai , Guowei Huang , Xingyue Quan , Hang Xu , Li Zhang

The 180x360 omnidirectional field of view captured by 360-degree cameras enables their use in a wide range of applications such as embodied AI and virtual reality. Although recent advances in multimodal large language models (MLLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Zihao Dongfang , Xu Zheng , Ziqiao Weng , Yuanhuiyi Lyu , Danda Pani Paudel , Luc Van Gool , Kailun Yang , Xuming Hu

Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering.…

Computation and Language · Computer Science 2025-11-11 Akshar Tumu , Varad Shinde , Parisa Kordjamshidi

Multimodal Large Language Models (MLLMs) have achieved impressive results on vision-language benchmarks, yet it remains unclear whether these benchmarks assess genuine global reasoning or allow success via localized visual cues. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Amit Agarwal , Hitesh Laxmichand Patel , Srikant Panda , Hansa Meghwani , Jyotika Singh , Karan Dua , Paul Li , Tao Sheng , Sujith Ravi , Dan Roth

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Diankun Wu , Fangfu Liu , Yi-Hsin Hung , Yueqi Duan

Despite recent advances demonstrating vision-language models' (VLMs) abilities to describe complex relationships in images using natural language, their capability to quantitatively reason about object sizes and distances remains…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Yuan-Hong Liao , Rafid Mahmood , Sanja Fidler , David Acuna

Spatial intelligence (SI) represents a cognitive ability encompassing the visualization, manipulation, and reasoning about spatial relationships, underpinning disciplines from neuroscience to robotics. We introduce SITE, a benchmark dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Wenqi Wang , Reuben Tan , Pengyue Zhu , Jianwei Yang , Zhengyuan Yang , Lijuan Wang , Andrey Kolobov , Jianfeng Gao , Boqing Gong

Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where beliefs must be continuously revised from egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chih-Ting Liao , Xi Xiao , Chunlei Meng , Zhangquan Chen , Yitong Qiao , Weilin Zhou , Tianyang Wang , Xu Zheng , Xin Cao

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Xiaoyu Zhan , Wenxuan Huang , Hao Sun , Xinyu Fu , Changfeng Ma , Shaosheng Cao , Bohan Jia , Shaohui Lin , Zhenfei Yin , Lei Bai , Wanli Ouyang , Yuanqi Li , Jie Guo , Yanwen Guo

Spatial reasoning is a fundamental aspect of human intelligence. One key concept in spatial cognition is the Frame of Reference, which identifies the perspective of spatial expressions. Despite its significance, FoR has received limited…

Computation and Language · Computer Science 2025-11-25 Tanawan Premsri , Parisa Kordjamshidi

We argue that progress in true multimodal intelligence calls for a shift from reactive, task-driven systems and brute-force long context towards a broader paradigm of supersensing. We frame spatial supersensing as four stages beyond…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Shusheng Yang , Jihan Yang , Pinzhi Huang , Ellis Brown , Zihao Yang , Yue Yu , Shengbang Tong , Zihan Zheng , Yifan Xu , Muhan Wang , Daohan Lu , Rob Fergus , Yann LeCun , Li Fei-Fei , Saining Xie

Embodied agents are expected to assist humans by actively exploring unknown environments and reasoning about spatial contexts. When deployed in real life, agents often face sequential tasks where each new task follows the completion of the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Zhongyi Cai , Yi Du , Chen Wang , Yu Kong

Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Chentao Li , Zirui Gao , Mingze Gao , Yinglian Ren , Jianjiang Feng , Jie Zhou

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly focus on task-level evaluation and fail to provide…

Spatial reasoning, the ability to ground language in 3D understanding, remains a persistent challenge for Vision-Language Models (VLMs). We identify two fundamental bottlenecks: inadequate 3D understanding capabilities stemming from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Yejie Guo , Yunzhong Hou , Wufei Ma , Meng Tang , Ming-Hsuan Yang

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in various multimodal tasks. To pursue higher intelligence in space, MLLMs require integrating multiple spatial capabilities, even for handling simple and normal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Ziyang Gong , Wenhao Li , Oliver Ma , Songyuan Li , Zhaokai Wang , Songyuan Li , Jiayi Ji , Xue Yang , Gen Luo , Junchi Yan , Rongrong Ji
‹ Prev 1 4 5 6 7 8 10 Next ›