English
Related papers

Related papers: Sparkle: Mastering Basic Spatial Capabilities in V…

200 papers

Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs'…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Weichen Zhang , Zile Zhou , Xin Zeng , Xuchen Liu , Jianjie Fang , Chen Gao , Yong Li , Jinqiang Cui , Xinlei Chen , Xiao-Ping Zhang

For tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense) knowledge. A representative task is Visual Question…

Computer Vision and Pattern Recognition · Computer Science 2018-12-12 Somak Aditya , Rudra Saha , Yezhou Yang , Chitta Baral

Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spatially coherent, plausibly executable actions in 3D digital…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Niyati Rawal , Sushant Ravva , Shah Alam Abir , Saksham Jain , Aman Chadha , Vinija Jain , Suranjana Trivedy , Amitava Das

Current vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessary for human-like understanding and real-world applications.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Wenyu Zhang , Wei En Ng , Lixin Ma , Yuwen Wang , Junqi Zhao , Allison Koenecke , Boyang Li , Lu Wang

Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leaving the inference-time approach relatively underexplored. In…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tingshu Mou , Jiabo He , Renying Wang , Ce Liu , Hao Yang , Tiehua Zhang , Jingjing Chen , Xingjun Ma

While current multimodal models can answer questions based on 2D images, they lack intrinsic 3D object perception, limiting their ability to comprehend spatial relationships and depth cues in 3D scenes. In this work, we propose N3D-VLM, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Yuxin Wang , Lei Ke , Boqiang Zhang , Tianyuan Qu , Hanxun Yu , Zhenpeng Huang , Meng Yu , Dan Xu , Dong Yu

Space grounding refers to localizing a set of spatial references described in natural language instructions. Traditional methods often fail to account for complex reasoning -- such as distance, geometry, and inter-object relationships --…

Robotics · Computer Science 2025-11-20 Nayoung Oh , Dohyun Kim , Junhyeong Bang , Rohan Paul , Daehyung Park

Vision-Language Models (VLMs) have emerged as general purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, also lacking some basic visual…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Shivam Chandhok , Wan-Cyuan Fan , Leonid Sigal

With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Peiwen Sun , Shiqiang Lang , Dongming Wu , Yi Ding , Kaituo Feng , Huadai Liu , Zhen Ye , Rui Liu , Yun-Hui Liu , Jianan Wang , Xiangyu Yue

Service robots are expected to reliably make sense of complex, fast-changing environments. From a cognitive standpoint, they need the appropriate reasoning capabilities and background knowledge required to exhibit human-like Visual…

Artificial Intelligence · Computer Science 2021-04-02 Agnese Chiatti , Gianluca Bardaro , Enrico Motta , Enrico Daga

Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Cheolhong Min , Jaeyun Jung , Daeun Lee , Hyeonseong Jeon , Yu Su , Jonathan Tremblay , Chan Hee Song , Jaesik Park

While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typically suffer from exorbitant modality-alignment costs and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yi Zhang , Youya Xia , Yong Wang , Meng Song , Xin Wu , Wenjun Wan , Bingbing Liu , AiXue Ye , Hongbo Zhang , Feng Wen

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems primarily from 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Kun Ouyang , Yuanxin Liu , Haoning Wu , Yi Liu , Hao Zhou , Jie Zhou , Fandong Meng , Xu Sun

Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks. However, the spatial intelligence capabilities of existing VLMs in UAV…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Lingfeng Zhang , Yuchen Zhang , Hongsheng Li , Haoxiang Fu , Yingbo Tang , Hangjun Ye , Long Chen , Xiaojun Liang , Xiaoshuai Hao , Wenbo Ding

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yolo Y. Tang , Pinxin Liu , Zhangyun Tan , Mingqian Feng , Rui Mao , Chao Huang , Jing Bi , Yunzhong Xiao , Susan Liang , Hang Hua , Ali Vosoughi , Luchuan Song , Zeliang Zhang , Chenliang Xu

We propose RocketScience, an open-source contrastive VLM benchmark that tests for spatial relation understanding. It is comprised of entirely new real-world image-text pairs covering mostly relative spatial understanding and the order of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Nils Hoehing , Mayug Maniparambil , Ellen Rushe , Noel E. O'Connor , Anthony Ventresque

Vision-language models (VLMs) achieve strong benchmark results, yet can exhibit systematic perceptual weaknesses: structured, large changes to pixel values can cause confident yet nonsensical predictions, even when the underlying scene…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Nicoleta-Nina Basoc , Adrian Cosma , Emilian Radoi

Recent progress in large language models (LLMs) has shown that reasoning improves when intermediate thoughts are externalized into explicit workspaces, such as chain-of-thought traces or tool-augmented reasoning. Yet, visual language models…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Oindrila Saha , Vojtech Krs , Radomir Mech , Subhransu Maji , Matheus Gadelha , Kevin Blackburn-Matzen

This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical orientations and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Jason Qiu , Zachary Meurer , Xavier Thomas , Deepti Ghadiyaram

Vision Language Models (VLMs) demonstrate remarkable proficiency in addressing a wide array of visual questions, which requires strong perception and reasoning faculties. Assessing these two competencies independently is crucial for model…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Yuxuan Qiao , Haodong Duan , Xinyu Fang , Junming Yang , Lin Chen , Songyang Zhang , Jiaqi Wang , Dahua Lin , Kai Chen