English
Related papers

Related papers: Boosting MLLM Spatial Reasoning with Geometrically…

200 papers

We introduce GSU, a text-only grid dataset to evaluate the spatial reasoning capabilities of LLMs over 3 core tasks: navigation, object localization, and structure composition. By forgoing visual inputs, isolating spatial reasoning from…

Computation and Language · Computer Science 2026-03-19 Risham Sidhu , Julia Hockenmaier

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even…

Computation and Language · Computer Science 2026-03-26 Xianzheng Ma , Tao Sun , Shuai Chen , Yash Bhalgat , Jindong Gu , Angel X Chang , Iro Armeni , Iro Laina , Songyou Peng , Victor Adrian Prisacariu

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Tanner Muturi , Blessing Agyei Kyem , Joshua Kofi Asamoah , Neema Jakisa Owor , Richard Dyzinela , Andrews Danyo , Yaw Adu-Gyamfi , Armstrong Aboah

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zhangyang Qi , Zhixiong Zhang , Ye Fang , Jiaqi Wang , Hengshuang Zhao

With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This convergence enables…

Robotics · Computer Science 2025-11-19 Vinit Mehta , Charu Sharma , Karthick Thiyagarajan

Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perception mechanisms, prior studies attempt to construct a spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yibin Huang , Wang Xu , Wanyue Zhang , Helu Zhi , Jingjing Huang , Yangbin Xu , Yangang Sun , Conghui Zhu , Tiejun Zhao

Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, particularly under large scale variation and ambiguous object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yanchun Cheng , Rundong Wang , Xulei Yang , Alok Prakash , Daniela Rus , Marcelo H Ang , ShiJie Li

The recent development of Large Language Models (LLMs) with strong reasoning ability has driven research in various domains such as mathematics, coding, and scientific discovery. Meanwhile, 3D visual grounding, as a fundamental task in 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Hsiang-Wei Huang , Kuang-Ming Chen , Wenhao Chai , Cheng-Yen Yang , Jen-Hao Cheng , Jenq-Neng Hwang

Large Language Models(LLMs) have revolutionized text generation and multimodal perception,but their capabilities in 3D content generation remain underexplored. Existing methods compromise by producing either low-resolution meshes or coarse…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Junming Huang , Chi Wang , Letian Li , Guangkai Xu , Donglin Huang , Hao Chen , Qiang Dai , Weiwei Xu

Large vision-language models (VLMs) show strong multimodal understanding but still struggle with 3D spatial reasoning, such as distance estimation, size comparison, and cross-view consistency. Existing 3D-aware methods either depend on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Ruosen Zhao , Zhikang Zhang , Jialei Xu , Jiahao Chang , Dong Chen , Lingyun Li , Weijian Sun , Zizhuang Wei

Developing a multi-modal language model capable of understanding 3D scenes remains challenging due to the limited availability of 3D training data, in contrast to the abundance of 2D datasets used for vision-language models (VLM). As an…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Doriand Petit , Steve Bourgeois , Vincent Gay-Bellile , Florian Chabot , Loïc Barthe

Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zhenyang Liu , Yikai Wang , Sixiao Zheng , Tongying Pan , Longfei Liang , Yanwei Fu , Xiangyang Xue

Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language…

Robotics · Computer Science 2025-11-03 Simindokht Jahangard , Mehrzad Mohammadi , Abhinav Dhall , Hamid Rezatofighi

The rising importance of 3D understanding, pivotal in computer vision, autonomous driving, and robotics, is evident. However, a prevailing trend, which straightforwardly resorted to transferring 2D alignment strategies to the 3D domain,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Jiayi Ji , Haowei Wang , Changli Wu , Yiwei Ma , Xiaoshuai Sun , Rongrong Ji

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross-modal integration. However, current benchmarks…

Computation and Language · Computer Science 2026-03-04 Shunki Uebayashi , Kento Masui , Kyohei Atarashi , Han Bao , Hisashi Kashima , Naoto Inoue , Mayu Otani , Koh Takeuchi

Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limited availability of 3D scene-language paired data, it is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Shengli Zhou , Minghang Zheng , Feng Zheng , Yang Liu

Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments. Despite recent progress, existing studies reveal that MLLMs still struggle with spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Wanyue Zhang , Yibin Huang , Yangbin Xu , JingJing Huang , Helu Zhi , Shuo Ren , Wang Xu , Jiajun Zhang

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang

Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still face two fundamental limitations. First, they struggle with coarse semantic alignment, often…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Weimin Bai , Yubo Li , Weijian Luo , Zeqiang Lai , Yequan Wang , Wenzheng Chen , He Sun

Large language models (LLMs) have been effectively used for many computer vision tasks, including image classification. In this paper, we present a simple yet effective approach for zero-shot image classification using multimodal LLMs.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Abdelrahman Abdelhamed , Mahmoud Afifi , Alec Go