English
Related papers

Related papers: RoboBrain 2.5: Depth in Sight, Time in Mind

200 papers

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 An-Chieh Cheng , Hongxu Yin , Yang Fu , Qiushan Guo , Ruihan Yang , Jan Kautz , Xiaolong Wang , Sifei Liu

This study investigates how adequate coordination among the different cognitive processes of a humanoid robot can be developed through end-to-end learning of direct perception of visuomotor stream. We propose a deep dynamic neural network…

Artificial Intelligence · Computer Science 2017-06-09 Jungsik Hwang , Jun Tani

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer that processes images at their native, variable…

Embodied intelligence fundamentally requires a capability to determine where to act in 3D space. We formalize this requirement as embodied localization -- the problem of predicting executable 3D points conditioned on visual observations and…

Robotics · Computer Science 2026-03-31 Qiming Zhu , Zhirui Fang , Tianming Zhang , Chuanxiu Liu , Xiaoke Jiang , Lei Zhang

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Yu Shang , Xin Zhang , Yinzhou Tang , Lei Jin , Chen Gao , Wei Wu , Yong Li

Embodied navigation requires agents to integrate perception, reasoning, and action for robust interaction in complex 3D environments. Existing approaches often suffer from incoherent and unstable reasoning traces that hinder generalization…

Robotics · Computer Science 2025-09-16 Qingxiang Liu , Ting Huang , Zeyu Zhang , Hao Tang

Autonomous service robots require computational frameworks that allow them to generalize knowledge to new situations in a manner that models uncertainty while scaling to real-world problem sizes. The Robot Common Sense Embedding (RoboCSE)…

Robotics · Computer Science 2019-03-04 Angel Daruna , Weiyu Liu , Zsolt Kira , Sonia Chernova

A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate grounding model. Incompatible decoders and box heads hinder…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Yani Zhang , Dongming Wu , Hao Shi , Yingfei Liu , Tiancai Wang , Xingping Dong

In this work, we investigate how spatially grounded auxiliary representations can provide both broad, high-level grounding as well as direct, actionable information to improve policy learning performance and generalization for dexterous…

Robotics · Computer Science 2025-06-09 Jonathan Yang , Chuyuan Kelly Fu , Dhruv Shah , Dorsa Sadigh , Fei Xia , Tingnan Zhang

Deep learning's success in perception, natural language processing, etc. inspires hopes for advancements in autonomous robotics. However, real-world robotics face challenges like variability, high-dimensional state spaces, non-linear…

Robotics · Computer Science 2025-01-28 Sven Behnke

Many language-guided robotic systems rely on collapsing spatial reasoning into discrete points, making them brittle to perceptual noise and semantic ambiguity. To address this challenge, we propose RoboMAP, a framework that represents…

Robotics · Computer Science 2025-10-16 Xinyu Shao , Yanzhe Tang , Pengwei Xie , Kaiwen Zhou , Yuzheng Zhuang , Xingyue Quan , Jianye Hao , Long Zeng , Xiu Li

Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs) excel at static perception tasks, they struggle with the…

Artificial Intelligence · Computer Science 2025-07-15 Di Wu , Jiaxin Fan , Junzhe Zang , Guanbo Wang , Wei Yin , Wenhao Li , Bo Jin

We propose a novel Transformer-based architecture for the task of generative modelling of 3D human motion. Previous work commonly relies on RNN-based models considering shorter forecast horizons reaching a stationary and often implausible…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Emre Aksan , Manuel Kaufmann , Peng Cao , Otmar Hilliges

Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under sensor noise, environmental variation, and platform shifts.…

Robotics · Computer Science 2026-01-09 Lingdong Kong , Shaoyuan Xie , Zeying Gong , Ye Li , Meng Chu , Ao Liang , Yuhao Dong , Tianshuai Hu , Ronghe Qiu , Rong Li , Hanjiang Hu , Dongyue Lu , Wei Yin , Wenhao Ding , Linfeng Li , Hang Song , Wenwei Zhang , Yuexin Ma , Junwei Liang , Zhedong Zheng , Lai Xing Ng , Benoit R. Cottereau , Wei Tsang Ooi , Ziwei Liu , Zhanpeng Zhang , Weichao Qiu , Wei Zhang , Ji Ao , Jiangpeng Zheng , Siyu Wang , Guang Yang , Zihao Zhang , Yu Zhong , Enzhu Gao , Xinhan Zheng , Xueting Wang , Shouming Li , Yunkai Gao , Siming Lan , Mingfei Han , Xing Hu , Dusan Malic , Christian Fruhwirth-Reisinger , Alexander Prutsch , Wei Lin , Samuel Schulter , Horst Possegger , Linfeng Li , Jian Zhao , Zepeng Yang , Yuhang Song , Bojun Lin , Tianle Zhang , Yuchen Yuan , Chi Zhang , Xuelong Li , Youngseok Kim , Sihwan Hwang , Hyeonjun Jeong , Aodi Wu , Xubo Luo , Erjia Xiao , Lingfeng Zhang , Yingbo Tang , Hao Cheng , Renjing Xu , Wenbo Ding , Lei Zhou , Long Chen , Hangjun Ye , Xiaoshuai Hao , Shuangzhi Li , Junlong Shen , Xingyu Li , Hao Ruan , Jinliang Lin , Zhiming Luo , Yu Zang , Cheng Wang , Hanshi Wang , Xijie Gong , Yixiang Yang , Qianli Ma , Zhipeng Zhang , Wenxiang Shi , Jingmeng Zhou , Weijun Zeng , Kexin Xu , Yuchen Zhang , Haoxiang Fu , Ruibin Hu , Yanbiao Ma , Xiyan Feng , Wenbo Zhang , Lu Zhang , Yunzhi Zhuge , Huchuan Lu , You He , Seungjun Yu , Junsung Park , Youngsun Lim , Hyunjung Shim , Faduo Liang , Zihang Wang , Yiming Peng , Guanyu Zong , Xu Li , Binghao Wang , Hao Wei , Yongxin Ma , Yunke Shi , Shuaipeng Liu , Dong Kong , Yongchun Lin , Huitong Yang , Liang Lei , Haoang Li , Xinliang Zhang , Zhiyong Wang , Xiaofeng Wang , Yuxia Fu , Yadan Luo , Djamahl Etchegaray , Yang Li , Congfei Li , Yuxiang Sun , Wenkai Zhu , Wang Xu , Linru Li , Longjie Liao , Jun Yan , Benwu Wang , Xueliang Ren , Xiaoyu Yue , Jixian Zheng , Jinfeng Wu , Shurui Qin , Wei Cong , Yao He

Depth sensors are widely deployed across robotic platforms, and advances in fast, high-fidelity depth simulation have enabled robotic policies trained on depth observations to achieve robust sim-to-real transfer for a wide range of tasks.…

Robotics · Computer Science 2026-01-28 Manthan Patel , Jonas Frey , Mayank Mittal , Fan Yang , Alexander Hansson , Amir Bar , Cesar Cadena , Marco Hutter

Functional magnetic resonance imaging produces high dimensional data, with a less then ideal number of labelled samples for brain decoding tasks (predicting brain states). In this study, we propose a new deep temporal convolutional neural…

Machine Learning · Computer Science 2015-01-13 Orhan Firat , Emre Aksan , Ilke Oztekin , Fatos T. Yarman Vural

We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and…

Sound · Computer Science 2024-01-17 Ashutosh Pandey , Buye Xu

Reasoning is central to purposeful action, yet most robotic foundation models map perception and instructions directly to control, which limits adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models…

While contemporary Vision-Language Models (VLMs) excel at 2D visual understanding, they remain constrained by a passive, 2D-centric paradigm that severely limits genuine 3D spatial reasoning. To bridge this gap, we introduce Think3D, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zaibin Zhang , Yuhan Wu , Lianjie Jia , Yifan Wang , Zhongbo Zhang , Yijiang Li , Binghao Ran , Fuxi Zhang , Zhuohan Sun , Zhenfei Yin , Lijun Wang , Huchuan Lu

Vision-centric hierarchical embodied models have demonstrated strong potential. However, existing methods lack spatial awareness capabilities, limiting their effectiveness in bridging visual plans to actionable control in complex…

Robotics · Computer Science 2025-11-19 Yijun Liu , Yuwei Liu , Yuan Meng , Jieheng Zhang , Yuwei Zhou , Ye Li , Jiacheng Jiang , Kangye Ji , Shijia Ge , Zhi Wang , Wenwu Zhu