English
Related papers

Related papers: Spatial Reasoning with Vision-Language Models in E…

200 papers

We introduce STSBench, a scenario-based framework to benchmark the holistic understanding of vision-language models (VLMs) for autonomous driving. The framework automatically mines pre-defined traffic scenarios from any dataset using…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Christian Fruhwirth-Reisinger , Dušan Malić , Wei Lin , David Schinagl , Samuel Schulter , Horst Possegger

The 180x360 omnidirectional field of view captured by 360-degree cameras enables their use in a wide range of applications such as embodied AI and virtual reality. Although recent advances in multimodal large language models (MLLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Zihao Dongfang , Xu Zheng , Ziqiao Weng , Yuanhuiyi Lyu , Danda Pani Paudel , Luc Van Gool , Kailun Yang , Xuming Hu

Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily emphasize indoor or street settings, overlooking the unique…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Haotian Xu , Yue Hu , Zhengqiu Zhu , Chen Gao , Ziyou Wang , Junreng Rao , Wenhao Lu , Weishi Li , Quanjun Yin , Yong Li

Reasoning about spatial relationships between objects is essential for many real-world robotic tasks, such as fetch-and-delivery, object rearrangement, and object search. The ability to detect and disambiguate different objects and identify…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Negar Nejatishahidin , Madhukar Reddy Vongala , Jana Kosecka

Existing Vision Language Models (VLMs) architecturally rooted in "flatland" perception, fundamentally struggle to comprehend real-world 3D spatial intelligence. This failure stems from a dual-bottleneck: input-stage conflict between…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Zhongbin Guo , Jiahe Liu , Yushan Li , Wenyu Gao , Zhen Yang , Chenzhi Li , Xinyue Zhang , Ping Jian

Recent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yujie Lu , Dongfu Jiang , Wenhu Chen , William Yang Wang , Yejin Choi , Bill Yuchen Lin

The spatial reasoning task aims to reason about the spatial relationships in 2D and 3D space, which is a fundamental capability for Visual Question Answering (VQA) and robotics. Although vision language models (VLMs) have developed rapidly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xun Liang , Xin Guo , Zhongming Jin , Weihang Pan , Penghui Shang , Deng Cai , Binbin Lin , Jieping Ye

Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spatially coherent, plausibly executable actions in 3D digital…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Niyati Rawal , Sushant Ravva , Shah Alam Abir , Saksham Jain , Aman Chadha , Vinija Jain , Suranjana Trivedy , Amitava Das

Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in Zero-Shot Egocentric Action Recognition…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Guangzhao Dai , Xiangbo Shu , Wenhao Wu , Rui Yan , Jiachao Zhang

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3D-Bench, a holistic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Tianxu Wang , Zhuofan Zhang , Ziyu Zhu , Yue Fan , Jing Xiong , Pengxiang Li , Xiaojian Ma , Qing Li

Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Mengdi Jia , Zekun Qi , Shaochen Zhang , Wenyao Zhang , Xinqiang Yu , Jiawei He , He Wang , Li Yi

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Muhammad Sohail Danish , Muhammad Akhtar Munir , Syed Roshaan Ali Shah , Kartik Kuckreja , Fahad Shahbaz Khan , Paolo Fraccaro , Alexandre Lacoste , Salman Khan

The advent of Multimodal Large Language Models, leveraging the power of Large Language Models, has recently demonstrated superior multimodal understanding and reasoning abilities, heralding a new era for artificial general intelligence.…

Artificial Intelligence · Computer Science 2025-04-14 Lu Qiu , Yi Chen , Yuying Ge , Yixiao Ge , Ying Shan , Xihui Liu

Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yogesh Kulkarni , Pooyan Fazli

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Yifan Liu , Fangneng Zhan , Kaichen Zhou , Yilun Du , Paul Pu Liang , Hanspeter Pfister

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high level relationships ('left…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Vineet Bhat , Sungsu Kim , Valts Blukis , Greg Heinrich , Prashanth Krishnamurthy , Ramesh Karri , Stan Birchfield , Farshad Khorrami , Jonathan Tremblay

Open-set perception in complex traffic environments poses a critical challenge for autonomous driving systems, particularly in identifying previously unseen object categories, which is vital for ensuring safety. Visual Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Fuhao Chang , Shuxin Li , Yabei Li , Lei He

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Dongchuan Ran , Linyu Ou , Xueheng Li , Wenwen Tong , Chenxu Guo , Hewei Guo , Kaibing Wang , Lewei Lu

Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. However, VLMs still exhibit several limitations in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Simone Alghisi , Gabriel Roccabruna , Massimo Rizzoli , Seyed Mahed Mousavi , Giuseppe Riccardi

Pre-trained general-purpose Vision-Language Models (VLM) hold the potential to enhance intuitive human-machine interactions due to their rich world knowledge and 2D object detection capabilities. However, VLMs for 3D coordinates detection…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ari Wahl , Dorian Gawlinski , David Przewozny , Paul Chojecki , Felix Bießmann , Sebastian Bosse