English
Related papers

Related papers: ESI-Bench: Towards Embodied Spatial Intelligence t…

200 papers

Embodied Artificial Intelligence (Embodied AI) is crucial for achieving Artificial General Intelligence (AGI) and serves as a foundation for various applications (e.g., intelligent mechatronics systems, smart manufacturing) that bridge…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yang Liu , Weixing Chen , Yongjie Bai , Xiaodan Liang , Guanbin Li , Wen Gao , Liang Lin

Optimizing and refining action execution through exploration and interaction is a promising way for robotic manipulation. However, practical approaches to interaction-driven robotic learning are still underexplored, particularly for…

Robotics · Computer Science 2025-09-24 Yibo Peng , Jiahao Yang , Shenhao Yan , Ziyu Huang , Shuang Li , Shuguang Cui , Yiming Zhao , Yatong Han

Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR…

Humans possess a remarkable capacity for spatial cognition, allowing for self-localization even in novel or unfamiliar environments. While hippocampal neurons encoding position and orientation are well documented, the large-scale neural…

Neurons and Cognition · Quantitative Biology 2025-07-17 Weichen Dai , Yuxuan Huang , Li Zhu , Dongjun Liu , Yu Zhang , Qibin Zhao , Andrzej Cichocki , Fabio Babiloni , Ke Li , Jianyu Qiu , Gangyong Jia , Wanzeng Kong , Qing Wu

We aim for zero-shot localization and classification of human actions in video. Where traditional approaches rely on global attribute or object classification scores for their zero-shot knowledge transfer, our main contribution is a…

Computer Vision and Pattern Recognition · Computer Science 2017-12-14 Pascal Mettes , Cees G. M. Snoek

Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Wenbo Hu , Yining Hong , Yanjun Wang , Leison Gao , Zibu Wei , Xingcheng Yao , Nanyun Peng , Yonatan Bitton , Idan Szpektor , Kai-Wei Chang

Coordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large…

Artificial Intelligence · Computer Science 2026-01-23 Li Kang , Xiufeng Song , Heng Zhou , Yiran Qin , Jie Yang , Xiaohong Liu , Philip Torr , Lei Bai , Zhenfei Yin

Spatial reasoning is a core aspect of human intelligence that allows perception, inference and planning in 3D environments. However, current vision-language models (VLMs) struggle to maintain geometric coherence and cross-view consistency…

Artificial Intelligence · Computer Science 2025-12-03 Qiyao Xue , Weichen Liu , Shiqi Wang , Haoming Wang , Yuyang Wu , Wei Gao

We propose a general self-supervised learning approach for spatial perception tasks, such as estimating the pose of an object relative to the robot, from onboard sensor readings. The model is learned from training episodes, by relying on: a…

Robotics · Computer Science 2021-07-20 Mirko Nava , Antonio Paolillo , Jérôme Guzzi , Luca Maria Gambardella , Alessandro Giusti

Multimodal large language models (MLLMs) are proficient in perception and instruction-following, but they still struggle with spatial reasoning: the ability to mentally track and manipulate objects across multiple views and over time.…

Artificial Intelligence · Computer Science 2025-12-30 Ryan Spencer , Roey Yaari , Ritvik Vemavarapu , Joyce Yang , Steven Ngo , Utkarsh Sharma

In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settings, this task poses two key challenges: 1) Spatial understanding is challenging in complex 3D environments with diverse spatial…

Robotics · Computer Science 2026-05-21 Zhizhao Liang , Yi-Lin Wei , Xuhang Chen , Mu Lin , Yi-Xiang He , Zhexi Luo , Jun-Hui Liu , Kun-Yu Lin , Wei-Shi Zheng

In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmental constraints across city-scale outdoor domains. We introduce Sentinel Challenge, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Xiangye Lin , Hongxin Zhang , Ruxi Deng , Qinhong Zhou , Chuang Gan

We study the task of embodied visual active learning, where an agent is set to explore a 3d environment with the goal to acquire visual scene understanding by actively selecting views for which to request annotation. While accurate on some…

Computer Vision and Pattern Recognition · Computer Science 2020-12-18 David Nilsson , Aleksis Pirinen , Erik Gärtner , Cristian Sminchisescu

Recent advances in multimodal large language models (MLLMs) have shown remarkable capabilities in integrating vision and language for complex reasoning. While most existing benchmarks evaluate models under offline settings with a fixed set…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Jingli Lin , Chenming Zhu , Runsen Xu , Xiaohan Mao , Xihui Liu , Tai Wang , Jiangmiao Pang

Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous navigation. Unfortunately, existing benchmarks are inadequate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Xinmiao Huang , Qisong He , Zhenglin Huang , Boxuan Wang , Zhuoyun Li , Guangliang Cheng , Yi Dong , Xiaowei Huang

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

Artificial Intelligence · Computer Science 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu

In embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments. Previous algorithms primarily rely on point cloud, which, despite offering precise geometric…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Xuewu Lin , Tianwei Lin , Lichao Huang , Hongyu Xie , Zhizhong Su

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step…

Robotics · Computer Science 2024-04-16 Roberto Bigazzi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Tanner Muturi , Blessing Agyei Kyem , Joshua Kofi Asamoah , Neema Jakisa Owor , Richard Dyzinela , Andrews Danyo , Yaw Adu-Gyamfi , Armstrong Aboah

Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language…

Robotics · Computer Science 2025-11-03 Simindokht Jahangard , Mehrzad Mohammadi , Abhinav Dhall , Hamid Rezatofighi