English
Related papers

Related papers: RoboTracer: Mastering Spatial Trace with Reasoning…

200 papers

Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perception mechanisms, prior studies attempt to construct a spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yibin Huang , Wang Xu , Wanyue Zhang , Helu Zhi , Jingjing Huang , Yangbin Xu , Yangang Sun , Conghui Zhu , Tiejun Zhao

While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reasoning remains highly challenging. Recent efforts to introduce Chain-of-Thought (CoT) reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Lang Sun , Ronghao Fu , Zhuoran Duan , Haoran Liu , Xueyan Liu , Bo Yang

Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, particularly under large scale variation and ambiguous object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yanchun Cheng , Rundong Wang , Xulei Yang , Alok Prakash , Daniela Rus , Marcelo H Ang , ShiJie Li

Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification, often confusing physically distinct motions by relying on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Hang Wu , Yujun Cai , Zehao Li , Haonan Ge , Bowen Sun , Junsong Yuan , Yiwei Wang

Capturing spatial relationships from visual inputs is a cornerstone of human-like general intelligence. Several previous studies have tried to enhance the spatial awareness of Vision-Language Models (VLMs) by adding extra expert encoders,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Rui Yang , Ziyu Zhu , Yanwei Li , Jingjia Huang , Shen Yan , Siyuan Zhou , Zhe Liu , Xiangtai Li , Shuangye Li , Wenqian Wang , Yi Lin , Hengshuang Zhao

Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Turhan Can Kargin , Wojciech Jasiński , Adam Pardyl , Bartosz Zieliński , Marcin Przewięźlikowski

Existing robot policies predominantly adopt the task-centric approach, requiring end-to-end task data collection. This results in limited generalization to new tasks and difficulties in pinpointing errors within long-horizon, multi-stage…

Humans naturally possess the spatial reasoning ability to form and manipulate images and structures of objects in space. There is an increasing effort to endow Vision-Language Models (VLMs) with similar spatial reasoning capabilities.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jiahuan Zhang , Shunwen Bai , Tianheng Wang , Kaiwen Guo , Kai Han , Guozheng Rao , Kaicheng Yu

Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision language models based…

Artificial Intelligence · Computer Science 2025-10-23 Jinrui Liu , Bingyan Nie , Boyu Li , Yaran Chen , Yuze Wang , Shunsen He , Haoran Li

Spatial reasoning is an essential problem in embodied AI research. Efforts to enhance spatial reasoning abilities through supplementary spatial data and fine-tuning have proven limited and ineffective when addressing complex embodied tasks,…

Spatial cognition is essential for human intelligence, enabling problem-solving through visual simulations rather than solely relying on verbal reasoning. However, existing AI benchmarks primarily assess verbal reasoning, neglecting the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Linjie Li , Mahtab Bigverdi , Jiawei Gu , Zixian Ma , Yinuo Yang , Ziang Li , Yejin Choi , Ranjay Krishna

Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language…

Robotics · Computer Science 2025-11-03 Simindokht Jahangard , Mehrzad Mohammadi , Abhinav Dhall , Hamid Rezatofighi

Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their…

Computation and Language · Computer Science 2025-12-01 Philip Schroeder , Ondrej Biza , Thomas Weng , Hongyin Luo , James Glass

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot…

Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under sensor noise, environmental variation, and platform shifts.…

Robotics · Computer Science 2026-01-09 Lingdong Kong , Shaoyuan Xie , Zeying Gong , Ye Li , Meng Chu , Ao Liang , Yuhao Dong , Tianshuai Hu , Ronghe Qiu , Rong Li , Hanjiang Hu , Dongyue Lu , Wei Yin , Wenhao Ding , Linfeng Li , Hang Song , Wenwei Zhang , Yuexin Ma , Junwei Liang , Zhedong Zheng , Lai Xing Ng , Benoit R. Cottereau , Wei Tsang Ooi , Ziwei Liu , Zhanpeng Zhang , Weichao Qiu , Wei Zhang , Ji Ao , Jiangpeng Zheng , Siyu Wang , Guang Yang , Zihao Zhang , Yu Zhong , Enzhu Gao , Xinhan Zheng , Xueting Wang , Shouming Li , Yunkai Gao , Siming Lan , Mingfei Han , Xing Hu , Dusan Malic , Christian Fruhwirth-Reisinger , Alexander Prutsch , Wei Lin , Samuel Schulter , Horst Possegger , Linfeng Li , Jian Zhao , Zepeng Yang , Yuhang Song , Bojun Lin , Tianle Zhang , Yuchen Yuan , Chi Zhang , Xuelong Li , Youngseok Kim , Sihwan Hwang , Hyeonjun Jeong , Aodi Wu , Xubo Luo , Erjia Xiao , Lingfeng Zhang , Yingbo Tang , Hao Cheng , Renjing Xu , Wenbo Ding , Lei Zhou , Long Chen , Hangjun Ye , Xiaoshuai Hao , Shuangzhi Li , Junlong Shen , Xingyu Li , Hao Ruan , Jinliang Lin , Zhiming Luo , Yu Zang , Cheng Wang , Hanshi Wang , Xijie Gong , Yixiang Yang , Qianli Ma , Zhipeng Zhang , Wenxiang Shi , Jingmeng Zhou , Weijun Zeng , Kexin Xu , Yuchen Zhang , Haoxiang Fu , Ruibin Hu , Yanbiao Ma , Xiyan Feng , Wenbo Zhang , Lu Zhang , Yunzhi Zhuge , Huchuan Lu , You He , Seungjun Yu , Junsung Park , Youngsun Lim , Hyunjung Shim , Faduo Liang , Zihang Wang , Yiming Peng , Guanyu Zong , Xu Li , Binghao Wang , Hao Wei , Yongxin Ma , Yunke Shi , Shuaipeng Liu , Dong Kong , Yongchun Lin , Huitong Yang , Liang Lei , Haoang Li , Xinliang Zhang , Zhiyong Wang , Xiaofeng Wang , Yuxia Fu , Yadan Luo , Djamahl Etchegaray , Yang Li , Congfei Li , Yuxiang Sun , Wenkai Zhu , Wang Xu , Linru Li , Longjie Liao , Jun Yan , Benwu Wang , Xueliang Ren , Xiaoyu Yue , Jixian Zheng , Jinfeng Wu , Shurui Qin , Wei Cong , Yao He

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Junfei Wu , Jian Guan , Kaituo Feng , Qiang Liu , Shu Wu , Liang Wang , Wei Wu , Tieniu Tan

Space grounding refers to localizing a set of spatial references described in natural language instructions. Traditional methods often fail to account for complex reasoning -- such as distance, geometry, and inter-object relationships --…

Robotics · Computer Science 2025-11-20 Nayoung Oh , Dohyun Kim , Junhyeong Bang , Rohan Paul , Daehyung Park

We present AutoTraces, an autoregressive vision-language-trajectory model for robot trajectory forecasting in humam-populated environments, which harnesses the inherent reasoning capabilities of large language models (LLMs) to model complex…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Teng Wang , Yanting Lu , Ruize Wang

The evolution of Remote Sensing Vision-Language Models(RS-VLMs) emphasizes the importance of transitioning from perception-centric recognition toward high-level deductive reasoning to enhance cognitive reliability in complex spatial tasks.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Wenshuai Li , Xiantai Xiang , Zixiao Wen , Guangyao Zhou , Ben Niu , Feng Wang , Lijia Huang , Qiantong Wang , Yuxin Hu

Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant…

Computation and Language · Computer Science 2025-10-14 Shiqi Chen , Tongyao Zhu , Ruochen Zhou , Jinghan Zhang , Siyang Gao , Juan Carlos Niebles , Mor Geva , Junxian He , Jiajun Wu , Manling Li