English
Related papers

Related papers: Towards Physics-informed Spatial Intelligence with…

200 papers

Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and sensor noise, spatial-domain models struggle to disentangle…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Jun Ma , Zhenye Yang , Ruichen Zhou , Pei Zhang , Huan Li , Jinpeng Chen

Visual autoregressive (VAR) models generate images through next-scale prediction, naturally achieving coarse-to-fine, fast, high-fidelity synthesis mirroring human perception. In practice, this hierarchy can drift at inference time, as…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Youngwoo Shin , Jiwan Hur , Junmo Kim

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal,…

Artificial Intelligence · Computer Science 2025-05-21 Joel Currie , Gioele Migno , Enrico Piacenti , Maria Elena Giannaccini , Patric Bach , Davide De Tommaso , Agnieszka Wykowska

Streetscapes are an essential component of urban space. Their assessment is presently either limited to morphometric properties of their mass skeleton or requires labor-intensive qualitative evaluations of visually perceived qualities. This…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Joan Perez , Giovanni Fusco

Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolution. Existing methods predominantly rely on implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Chenjie Yang , Yutian Jiang , Yutong Deng , Chenyu Wu

Grid-centric perception is a crucial field for mobile robot perception and navigation. Nonetheless, grid-centric perception is less prevalent than object-centric perception as autonomous vehicles need to accurately perceive highly dynamic,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Yining Shi , Kun Jiang , Jiusi Li , Zelin Qian , Junze Wen , Mengmeng Yang , Ke Wang , Diange Yang

Perspective-Aware AI requires modeling evolving internal states--goals, emotions, contexts--not merely preferences. Progress is limited by a data bottleneck: digital footprints are privacy-sensitive and perspective states are rarely…

Artificial Intelligence · Computer Science 2026-02-17 Jisung Shin , Daniel Platnick , Marjan Alirezaie , Hossein Rahnama

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Wufei Ma , Luoxin Ye , Celso M de Melo , Jieneng Chen , Alan Yuille

Existing vision-language models often suffer from spatial hallucinations, i.e., generating incorrect descriptions about the relative positions of objects in an image. We argue that this problem mainly stems from the asymmetric properties…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Hang Yin , Xiaomin He , PeiWen Yuan , Yiwei Li , Jiayi Shi , Wenxiao Fan , Shaoxiong Feng , Kan Li

Large Language Models (LLMs), such as ChatGPT, demonstrate a strong understanding of human natural language and have been explored and applied in various fields, including reasoning, creative writing, code generation, translation, and…

Artificial Intelligence · Computer Science 2023-05-30 Zhenlong Li , Huan Ning

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining minimal manual supervision with large-scale Visual Question-Answering (VQA) pairs…

Robotics · Computer Science 2026-05-05 Yangzhe Kong , Daeun Song , Jing Liang , Dinesh Manocha , Ziyu Yao , Xuesu Xiao

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jingyu Li , Junjie Wu , Dongnan Hu , Xiangkai Huang , Bin Sun , Zhihui Hao , Xianpeng Lang , Xiatian Zhu , Li Zhang

Developing machine intelligence abilities in robots and autonomous systems is an expensive and time consuming process. Existing solutions are tailored to specific applications and are harder to generalize. Furthermore, scarcity of training…

Robotics · Computer Science 2023-10-10 Sai Vemprala , Shuhang Chen , Abhinav Shukla , Dinesh Narayanan , Ashish Kapoor

Humans excel at forming mental maps of their surroundings, equipping them to understand object relationships and navigate based on language queries. Our previous work, SI Maps (Nanwani L, Agarwal A, Jain K, et al. Instance-level semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Laksh Nanwani , Kumaraditya Gupta , Aditya Mathur , Swayam Agrawal , A. H. Abdul Hafez , K. Madhava Krishna

Existing driving style recognition systems largely depend on low-level sensor-derived features for training, neglecting the rich semantic reasoning capability inherent to human experts. This discrepancy results in a fundamental misalignment…

Robotics · Computer Science 2026-05-06 Zhaokun Chen , Chaopeng Zhang , Xiaohan Li , Wenshuo Wang , Gentiane Venture , Junqiang Xi

Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are too coarse to specify which regions actually support, contain, or contact one another, leading…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yinuo Bai , Peijun Xu , Kuixiang Shao , Yuyang Jiao , Jingxuan Zhang , Kaixin Yao , Jiayuan Gu , Jingyi Yu

Explainability is essential for autonomous vehicles and other robotics systems interacting with humans and other objects during operation. Humans need to understand and anticipate the actions taken by the machines for trustful and safe…

Artificial Intelligence · Computer Science 2024-07-09 Chen Tang , Nishan Srishankar , Sujitha Martin , Masayoshi Tomizuka

Spatial intelligence is central to embodied cognition, yet contemporary AI systems still struggle to reason about physical interactions in open-world human environments. Despite strong performance on controlled benchmarks, vision-language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Tianjun Gu , Jingyu Gong , Zhizhong Zhang , Yuan Xie , Lizhuang Ma , Xin Tan , Athanasios V

Multimodal models have achieved remarkable progress in recent years. Nevertheless, they continue to exhibit notable limitations in spatial understanding and reasoning, the very capability that anchors artificial general intelligence in the…