English
Related papers

Related papers: RoboSpatial: Teaching Spatial Understanding to 2D …

200 papers

A key challenge in robot manipulation lies in developing policy models with strong spatial understanding, the ability to reason about 3D geometry, object relations, and robot embodiment. Existing methods often fall short: 3D point cloud…

Robotics · Computer Science 2025-09-25 Xuewu Lin , Tianwei Lin , Lichao Huang , Hongyu Xie , Yiwei Jin , Keyu Li , Zhizhong Su

Genuine spatial reasoning relies on the capacity to construct and manipulate coherent internal spatial representations, often conceptualized as mental models, rather than merely processing surface linguistic associations. While large…

Artificial Intelligence · Computer Science 2026-03-04 Peiyao Jiang , Zequn Qin , Xi Li

Recent advances in imitation learning have shown significant promise for robotic control and embodied intelligence. However, achieving robust generalization across diverse mounted camera observations remains a critical challenge. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Travis Davies , Jiahuan Yan , Xiang Chen , Yu Tian , Yueting Zhuang , Yiqi Huang , Luhui Hu

Holistic scene understanding poses a fundamental contribution to the autonomous operation of a robotic agent in its environment. Key ingredients include a well-defined representation of the surroundings to capture its spatial structure as…

Robotics · Computer Science 2024-05-24 Niclas Vödisch

Spatial reasoning in large-scale 3D environments remains challenging for current vision-language models, which are typically constrained to room-scale scenarios. We introduce H$^2$U3D (Holistic House Understanding in 3D), a 3D visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Hongpei Zheng , Shijie Li , Yanran Li , Hujun Yin

Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their…

Robotics · Computer Science 2025-05-28 Yiqi Huang , Travis Davies , Jiahuan Yan , Jiankai Sun , Xiang Chen , Luhui Hu

Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Xingyu Miao , Haoran Duan , Quanhao Qian , Jiuniu Wang , Yang Long , Ling Shao , Deli Zhao , Ran Xu , Gongjie Zhang

We present a system for generating and understanding of dynamic and static spatial relations in robotic interaction setups. Robots describe an environment of moving blocks using English phrases that include spatial relations such as…

Artificial Intelligence · Computer Science 2016-07-21 Michael Spranger , Jakob Suchan , Mehul Bhatt

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we introduce Ego3D Position Encoding…

Robotics · Computer Science 2025-05-20 Delin Qu , Haoming Song , Qizhi Chen , Yuanqi Yao , Xinyi Ye , Yan Ding , Zhigang Wang , JiaYuan Gu , Bin Zhao , Dong Wang , Xuelong Li

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal,…

Artificial Intelligence · Computer Science 2025-05-21 Joel Currie , Gioele Migno , Enrico Piacenti , Maria Elena Giannaccini , Patric Bach , Davide De Tommaso , Agnieszka Wykowska

The imagination of the surrounding environment based on experience and semantic cognition has great potential to extend the limited observations and provide more information for mapping, collision avoidance, and path planning. This paper…

Robotics · Computer Science 2021-04-09 Zhengcheng Shen , Linh Kästner , Jens Lambrecht

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Diankun Wu , Fangfu Liu , Yi-Hsin Hung , Yueqi Duan

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 An-Chieh Cheng , Hongxu Yin , Yang Fu , Qiushan Guo , Ruihan Yang , Jan Kautz , Xiaolong Wang , Sifei Liu

Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we introduce SpatialViLT, an enhanced VLM that integrates…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Chashi Mahiul Islam , Oteo Mamo , Samuel Jacob Chacko , Xiuwen Liu , Weikuan Yu

Spatial understanding is a fundamental problem with wide-reaching real-world applications. The representation of spatial knowledge is often modeled with spatial templates, i.e., regions of acceptability of two objects under an explicit…

Artificial Intelligence · Computer Science 2020-03-09 Guillem Collell , Luc Van Gool , Marie-Francine Moens

Recognizing spatial relations and reasoning about them is essential in multiple applications including navigation, direction giving and human-computer interaction in general. Spatial relations between objects can either be explicit --…

Computation and Language · Computer Science 2020-07-21 Soham Dan , Hangfeng He , Dan Roth

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view…

Robotics · Computer Science 2026-03-05 Yu Sheng , Lidian Wang , Xiaomeng Chu , Jiajun Deng , Min Cheng , Yanyong Zhang , Bei Hua , Houqiang Li , Jianmin Ji

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional…

Mapping the surrounding environment is essential for the successful operation of autonomous robots. While extensive research has focused on mapping geometric structures and static objects, the environment is also influenced by the movement…

Robotics · Computer Science 2023-09-04 Junyi Shi , Tomasz Piotr Kucner

Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments. Despite recent progress, existing studies reveal that MLLMs still struggle with spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Wanyue Zhang , Yibin Huang , Yangbin Xu , JingJing Huang , Helu Zhi , Shuo Ren , Wang Xu , Jiajun Zhang