English
Related papers

Related papers: SpatialLLM: A Compound 3D-Informed Design towards …

200 papers

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks…

Artificial Intelligence · Computer Science 2026-05-08 Peiran Xu , Sudong Wang , Yao Zhu , Jianing Li , Gege Qi , Yunjian Zhang

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Zuntao Liu , Yi Du , Taimeng Fu , Shaoshu Su , Cherie Ho , Chen Wang

This paper introduces Scene-LLM, a 3D-visual-language model that enhances embodied agents' abilities in interactive 3D indoor environments by integrating the reasoning strengths of Large Language Models (LLMs). Scene-LLM adopts a hybrid 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Rao Fu , Jingyu Liu , Xilun Chen , Yixin Nie , Wenhan Xiong

Visual Language Models (VLMs) are essential for various tasks, particularly visual reasoning tasks, due to their robust multi-modal information integration, visual reasoning capabilities, and contextual awareness. However, existing \VLMs{}'…

Computation and Language · Computer Science 2024-09-13 Zaiqiao Meng , Hao Zhou , Yifang Chen

While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typically suffer from exorbitant modality-alignment costs and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yi Zhang , Youya Xia , Yong Wang , Meng Song , Xin Wu , Wenjun Wan , Bingbing Liu , AiXue Ye , Hongbo Zhang , Feng Wen

Computer-aided design (CAD) significantly enhances the efficiency, accuracy, and innovation of design processes by enabling precise 2D and 3D modeling, extensive analysis, and optimization. Existing methods for creating CAD models rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Siyu Wang , Cailian Chen , Xinyi Le , Qimin Xu , Lei Xu , Yanzhou Zhang , Jie Yang

Despite the significant advancements in natural language processing capabilities demonstrated by large language models such as ChatGPT, their proficiency in comprehending and processing spatial information, especially within the domains of…

Computation and Language · Computer Science 2023-12-07 He Yan , Xinyao Hu , Xiangpeng Wan , Chengyu Huang , Kai Zou , Shiqi Xu

Uncrewed Aerial Vehicles (UAVs) are widely deployed across diverse applications due to their mobility and agility. Recent advances in Large Language Models (LLMs) offer a transformative opportunity to enhance UAV intelligence beyond…

Robotics · Computer Science 2026-03-03 Yousef Emami , Hao Zhou , Radha Reddy , Atefeh Hajijamali Arani , Biliang Wang , Kai Li , Luis Almeida , Zhu Han

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Chan Hee Song , Valts Blukis , Jonathan Tremblay , Stephen Tyree , Yu Su , Stan Birchfield

Spatial intelligence is crucial for vision--language models (VLMs) in the physical world, yet many benchmarks evaluate largely unconstrained scenes where models can exploit 2D shortcuts. We introduce SSI-Bench, a VQA benchmark for spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Chen Yang , Guanxin Lin , Youquan He , Peiyao Chen , Guanghe Liu , Yufan Mo , Zhouyuan Xu , Linhao Wang , Guohui Zhang , Zihang Zhang , Shenxiang Zeng , Chen Wang , Jiansheng Fan

Despite significant recent progress of Multimodal Large Language Models (MLLMs), current MLLMs are challenged by "spatio-temporal" prompts, i.e., prompts that refer to 1) the entirety of an environment encoded in a point cloud that the MLLM…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Haozhen Zheng , Beitong Tian , Mingyuan Wu , Zhenggang Tang , Klara Nahrstedt , Alex Schwing

The potential for Large Language Models (LLMs) to generate new information offers a potential step change for research and innovation. This is challenging to assert as it can be difficult to determine what an LLM has previously seen during…

Computation and Language · Computer Science 2024-05-24 Thomas Greatrix , Roger Whitaker , Liam Turner , Walter Colombo

Recent advances in scene understanding have leveraged multimodal large language models (MLLMs) for 3D reasoning by capitalizing on their strong 2D pretraining. However, the lack of explicit 3D data during MLLM pretraining limits 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Xiaohu Huang , Jingjing Wu , Qunyi Xie , Kai Han

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a benchmark to evaluate whether video-large language models…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Baining Zhao , Jianjie Fang , Zichao Dai , Ziyou Wang , Jirong Zha , Weichen Zhang , Chen Gao , Yue Wang , Jinqiang Cui , Xinlei Chen , Yong Li

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason about space, they only…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Arijit Ray , Jiafei Duan , Ellis Brown , Reuben Tan , Dina Bashkirova , Rose Hendrix , Kiana Ehsani , Aniruddha Kembhavi , Bryan A. Plummer , Ranjay Krishna , Kuo-Hao Zeng , Kate Saenko

The spatial reasoning task aims to reason about the spatial relationships in 2D and 3D space, which is a fundamental capability for Visual Question Answering (VQA) and robotics. Although vision language models (VLMs) have developed rapidly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xun Liang , Xin Guo , Zhongming Jin , Weihang Pan , Penghui Shang , Deng Cai , Binbin Lin , Jieping Ye

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiahui Zhang , Yurui Chen , Yanpeng Zhou , Yueming Xu , Ze Huang , Jilin Mei , Junhui Chen , Yu-Jie Yuan , Xinyue Cai , Guowei Huang , Xingyue Quan , Hang Xu , Li Zhang

Vision language models (VLMs) perform well on many tasks but often fail at spatial reasoning, which is essential for navigation and interaction with physical environments. Many spatial reasoning tasks depend on fundamental two-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Yihong Tang , Ao Qu , Zhaokai Wang , Dingyi Zhuang , Zhaofeng Wu , Wei Ma , Shenhao Wang , Yunhan Zheng , Zhan Zhao , Jinhua Zhao

Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA) frameworks typically use 2D representations to connect…

Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Ilias Stogiannidis , Steven McDonagh , Sotirios A. Tsaftaris
‹ Prev 1 4 5 6 7 8 10 Next ›