中文
相关论文

相关论文: RoboBrain 2.5: Depth in Sight, Time in Mind

200 篇论文

We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through the task of converting apartment photographs into accurate 2D floor plans. While the input modality (photographs) is well…

人工智能 · 计算机科学 2025-10-01 Lukas Petersson , Axel Backlund , Axel Wennstöm , Hanna Petersson , Callum Sharrock , Arash Dabiri

Spatial visual perception is a fundamental requirement in physical-world applications like autonomous driving and robotic manipulation, driven by the need to interact with 3D environments. Capturing pixel-aligned metric depth using RGB-D…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Bin Tan , Changjiang Sun , Xiage Qin , Hanat Adai , Zelin Fu , Tianxiang Zhou , Han Zhang , Yinghao Xu , Xing Zhu , Yujun Shen , Nan Xue

Recent advancements in legged robot perceptive locomotion have shown promising progress. However, terrain-aware humanoid locomotion remains largely constrained to two paradigms: depth image-based end-to-end learning and elevation map-based…

机器人学 · 计算机科学 2025-10-13 Jingkai Sun , Gang Han , Pihai Sun , Wen Zhao , Jiahang Cao , Jiaxu Wang , Yijie Guo , Qiang Zhang

Achieving robust spatial reasoning remains a fundamental challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical shortcuts via 3D grounding data or remain confined to 2D visual perception,…

人工智能 · 计算机科学 2026-03-11 Shouwei Ruan , Bin Wang , Zhenyu Wu , Qihui Zhu , Yuxiang Zhang , Hang Su , Yubin Wang

Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing…

机器人学 · 计算机科学 2025-05-01 Haifeng Huang , Xinyi Chen , Yilun Chen , Hao Li , Xiaoshen Han , Zehan Wang , Tai Wang , Jiangmiao Pang , Zhou Zhao

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Wufei Ma , Haoyu Chen , Guofeng Zhang , Yu-Cheng Chou , Jieneng Chen , Celso M de Melo , Alan Yuille

As the basis for prehensile manipulation, it is vital to enable robots to grasp as robustly as humans. Our innate grasping system is prompt, accurate, flexible, and continuous across spatial and temporal domains. Few existing methods cover…

机器人学 · 计算机科学 2023-06-07 Hao-Shu Fang , Chenxi Wang , Hongjie Fang , Minghao Gou , Jirong Liu , Hengxu Yan , Wenhai Liu , Yichen Xie , Cewu Lu

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their…

机器人学 · 计算机科学 2024-08-12 Xincheng Pang , Wenke Xia , Zhigang Wang , Bin Zhao , Di Hu , Dong Wang , Xuelong Li

The primary aim of this manuscript is to underscore a significant limitation in current deep learning models, particularly vision models. Unlike human vision, which efficiently selects only the essential visual areas for further processing,…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Ali Borji

Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but they continue to struggle with spatial understanding. Existing spatial MLLMs often rely on explicit 3D inputs or architecture-specific…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Hunar Batra , Haoqin Tu , Hardy Chen , Yuanze Lin , Cihang Xie , Ronald Clark

We present ProtoDepth, a novel prototype-based approach for continual learning of unsupervised depth completion, the multimodal 3D reconstruction task of predicting dense depth maps from RGB images and sparse point clouds. The unsupervised…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Patrick Rim , Hyoungseob Park , S. Gangopadhyay , Ziyao Zeng , Younjoon Chung , Alex Wong

Building robots that can automate labor-intensive tasks has long been the core motivation behind the advancements in computer vision and the robotics community. Recent interest in leveraging 3D algorithms, particularly neural fields, has…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Litian Liang , Liuyu Bian , Caiwei Xiao , Jialin Zhang , Linghao Chen , Isabella Liu , Fanbo Xiang , Zhiao Huang , Hao Su

Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings. A key challenge in applying them to robotic tasks is the modality gap between visual data and…

机器人学 · 计算机科学 2024-10-18 Ruoshi Liu , Alper Canberk , Shuran Song , Carl Vondrick

Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Most existing works improve spatial reasoning by introducing…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Zhenghao Chen , Huiqun Wang , Di Huang

Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using…

机器人学 · 计算机科学 2026-01-19 Dongyoung Kim , Sumin Park , Huiwon Jang , Jinwoo Shin , Jaehyung Kim , Younggyo Seo

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack…

机器学习 · 计算机科学 2025-12-02 Jacob Thompson , Emiliano Garcia-Lopez , Yonatan Bisk

For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper,…

机器人学 · 计算机科学 2021-11-08 Junha Roh , Karthik Desingh , Ali Farhadi , Dieter Fox

Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal models. Among its core challenges, robot manipulation stands…

Tactile feedback is critical for understanding the dynamics of both rigid and deformable objects in many manipulation tasks, such as non-prehensile manipulation and dense packing. We introduce an approach that combines visual and tactile…

机器人学 · 计算机科学 2024-07-02 Bo Ai , Stephen Tian , Haochen Shi , Yixuan Wang , Cheston Tan , Yunzhu Li , Jiajun Wu

Enabling robots to execute long-horizon manipulation tasks from free-form language instructions remains a fundamental challenge in embodied AI. While vision-language models (VLMs) have shown promise as high-level planners, their deployment…

机器人学 · 计算机科学 2025-10-01 Zitong Bo , Yue Hu , Jinming Ma , Mingliang Zhou , Junhui Yin , Yachen Kang , Yuqi Liu , Tong Wu , Diyun Xiang , Hao Chen