English
Related papers

Related papers: SORT3D: Spatial Object-centric Reasoning Toolbox f…

200 papers

Recent advances in large multimodal models suggest that explicit reasoning mechanisms play a critical role in improving model reliability, interpretability, and cross-modal alignment. While such reasoning-centric approaches have been proven…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Tianjiao Yu , Xinzhuo Li , Yifan Shen , Yuanzhe Liu , Ismini Lourentzou

Recent advances in 3D scene-language understanding have leveraged Large Language Models (LLMs) for 3D reasoning by transferring their general reasoning ability to 3D multi-modal contexts. However, existing methods typically adopt standard…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yerim Jeon , Miso Lee , WonJun Moon , Jae-Pil Heo

For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper,…

Robotics · Computer Science 2021-11-08 Junha Roh , Karthik Desingh , Ali Farhadi , Dieter Fox

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on labeled 3D data and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Rong Li , Shijie Li , Lingdong Kong , Xulei Yang , Junwei Liang

Object navigation (ObjectNav) requires an agent to navigate through unseen environments to find queried objects. Many previous methods attempted to solve this task by relying on supervised or reinforcement learning, where they are trained…

Computation and Language · Computer Science 2024-03-26 Yuxuan Kuang , Hai Lin , Meng Jiang

Deployable service and delivery robots struggle to navigate multi-floor buildings to reach object goals, as existing systems fail due to single-floor assumptions and requirements for offline, globally consistent maps. Multi-floor…

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Tanner Muturi , Blessing Agyei Kyem , Joshua Kofi Asamoah , Neema Jakisa Owor , Richard Dyzinela , Andrews Danyo , Yaw Adu-Gyamfi , Armstrong Aboah

3D visual grounding (3DVG) aims to locate objects in a 3D scene with natural language descriptions. Supervised methods have achieved decent accuracy, but have a closed vocabulary and limited language understanding ability. Zero-shot methods…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Qihao Yuan , Kailai Li , Jiaming Zhang

Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promising results but rely on view-dependent reasoning or implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xuefei Sun , Xujia Zhang , Brendan Crowe , Doncey Albin , Christoffer Heckman

3D Visual Grounding (3DVG) aims to localize target objects within a 3D scene based on natural language queries. To alleviate the reliance on costly 3D training data, recent studies have explored zero-shot 3DVG by leveraging the extensive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Zhao Jin , Rong-Cheng Tu , Jingyi Liao , Wenhao Sun , Xiao Luo , Shunyu Liu , Dacheng Tao

Grounded understanding of natural language in physical scenes can greatly benefit robots that follow human instructions. In object manipulation scenarios, existing end-to-end models are proficient at understanding semantic concepts, but…

Robotics · Computer Science 2023-04-03 Qian Luo , Yunfei Li , Yi Wu

Home-assistant robots have been a long-standing research topic, and one of the biggest challenges is searching for required objects in housing environments. Previous object-goal navigation requires the robot to search for a target object…

Robotics · Computer Science 2023-03-14 Hongyi Chen , Ruinian Xu , Shuo Cheng , Patricio A. Vela , Danfei Xu

3D object detection is fundamental for spatial understanding. Real-world environments demand models capable of recognizing diverse, previously unseen objects, which remains a major limitation of closed-set methods. Existing open-vocabulary…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Andrey Lemeshko , Bulat Gabdullin , Nikita Drozdov , Anton Konushin , Danila Rukhovich , Maksim Kolodiazhnyi

The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Jiaxin Huang , Ziwen Li , Hanlve Zhang , Runnan Chen , Xiao He , Yandong Guo , Wenping Wang , Tongliang Liu , Mingming Gong

3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing methods rely on a pre-defined Object Lookup Table (OLT) to query Visual Language Models (VLMs) for reasoning about object locations,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Wenyuan Huang , Zhao Wang , Zhou Wei , Ting Huang , Fang Zhao , Jian Yang , Zhenyu Zhang

Spatial question answering over egocentric video is a challenging task that requires Vision-Language Models (VLMs) to reason about 3D object positions, scene affordances, and directional relationships, particularly in the zero-shot setting…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Pawat Chunhachatrachai , Gueter Josmy Faure , Hung-Ting Su , Winston H. Hsu

Abstract semantic 3D scene understanding is a problem of critical importance in robotics. As robots still lack the common-sense knowledge about household objects and locations of an average human, we investigate the use of pre-trained…

Robotics · Computer Science 2023-11-09 William Chen , Siyi Hu , Rajat Talak , Luca Carlone

As robots begin to cohabit with humans in semi-structured environments, the need arises to understand instructions involving rich variability---for instance, learning to ground symbols in the physical world. Realistically, this task must…

Artificial Intelligence · Computer Science 2017-06-02 Yordan Hristov , Svetlin Penkov , Alex Lascarides , Subramanian Ramamoorthy

Precise spatial modeling in the operating room (OR) is foundational to many clinical tasks, supporting intraoperative awareness, hazard avoidance, and surgical decision-making. While existing approaches leverage large-scale multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Peiqi He , Zhenhao Zhang , Yixiang Zhang , Xiongjun Zhao , Shaoliang Peng

3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary, which can be restrictive. To address this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zhihao Yuan , Jinke Ren , Chun-Mei Feng , Hengshuang Zhao , Shuguang Cui , Zhen Li