中文
相关论文

相关论文: SAGA: Open-World Mobile Manipulation via Structure…

200 篇论文

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive…

机器人学 · 计算机科学 2026-05-22 Pengteng Li , Weiyu Guo , He Zhang , Tiefu Cai , Xiao He , Yandong Guo , Hui Xiong

Flexible, goal-directed behavior is a fundamental aspect of human life. Based on the free energy minimization principle, the theory of active inference formalizes the generation of such behavior from a computational neuroscience…

人工智能 · 计算机科学 2022-08-03 Fedor Scholz , Christian Gumbsch , Sebastian Otte , Martin V. Butz

Dexterous manipulation requires precise geometric reasoning, yet existing visuo-tactile learning methods struggle with sub-millimeter precision tasks that are routine for traditional model-based approaches. We identify a key limitation:…

机器人学 · 计算机科学 2026-02-27 Jialei Huang , Yang Ye , Yuanqing Gong , Xuezhou Zhu , Yang Gao , Kaifeng Zhang

To perform versatile mobile manipulation tasks in human-centered environments, the ability to efficiently transfer learned tasks and experiences from one robot to another or across different environments is key. In this paper, we present…

机器人学 · 计算机科学 2024-03-22 Christoph Pohl , Fabian Reister , Fabian Peller-Konrad , Tamim Asfour

Planning in realistic environments requires searching in large planning spaces. Affordances are a powerful concept to simplify this search, because they model what actions can be successful in a given situation. However, the classical…

机器人学 · 计算机科学 2021-06-24 Danfei Xu , Ajay Mandlekar , Roberto Martín-Martín , Yuke Zhu , Silvio Savarese , Li Fei-Fei

The utilization of broad datasets has proven to be crucial for generalization for a wide range of fields. However, how to effectively make use of diverse multi-task data for novel downstream tasks still remains a grand challenge in…

机器人学 · 计算机科学 2023-04-19 Kuan Fang , Patrick Yin , Ashvin Nair , Homer Walke , Gengchen Yan , Sergey Levine

This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Jiazhong Cen , Jiemin Fang , Chen Yang , Lingxi Xie , Xiaopeng Zhang , Wei Shen , Qi Tian

This paper presents a Surface-Aligned Gaussian representation for creating animatable human avatars from monocular videos,aiming at improving the novel view and pose synthesis performance while ensuring fast training and real-time…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Ronghan Chen , Yang Cong , Jiayue Liu

Reasoning about object affordances allows an autonomous agent to perform generalised manipulation tasks among object instances. While current approaches to grasp affordance estimation are effective, they are limited to a single hypothesis.…

机器人学 · 计算机科学 2019-06-25 Paola Ardón , Èric Pairet , Ronald P. A. Petrick , Subramanian Ramamoorthy , Katrin S. Lohan

Affordances have been introduced in literature as action opportunities that objects offer, and used in robotics to semantically represent their interconnection. However, when considering an environment instead of an object, the problem…

机器人学 · 计算机科学 2016-07-04 Francesco Riccio , Roberto Capobianco , Marc Hanheide , Daniele Nardi

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines,…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Hongchi Xia , Xuan Li , Zhaoshuo Li , Qianli Ma , Jiashu Xu , Ming-Yu Liu , Yin Cui , Tsung-Yi Lin , Wei-Chiu Ma , Shenlong Wang , Shuran Song , Fangyin Wei

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in…

Humans perceive and interact with the world with the awareness of equivariance, facilitating us in manipulating different objects in diverse poses. For robotic manipulation, such equivariance also exists in many scenarios. For example, no…

机器人学 · 计算机科学 2024-08-08 Yue Chen , Chenrui Tie , Ruihai Wu , Hao Dong

At its core, robotic manipulation is a problem of vision-to-geometry mapping ($f(v) \rightarrow G$). Physical actions are fundamentally defined by geometric properties like 3D positions and spatial relationships. Consequently, we argue that…

机器人学 · 计算机科学 2026-04-15 Zijian Song , Qichang Li , Jiawei Zhou , Zhenlong Yuan , Tianshui Chen , Liang Lin , Guangrun Wang

We introduce an open-source system called SIGMA (short for "Situated Interactive Guidance, Monitoring, and Assistance") as a platform for conducting research on task-assistive agents in mixed-reality scenarios. The system leverages the…

人机交互 · 计算机科学 2024-05-24 Dan Bohus , Sean Andrist , Nick Saw , Ann Paradiso , Ishani Chakraborty , Mahdi Rad

In order to enable robust operation in unstructured environments, robots should be able to generalize manipulation actions to novel object instances. For example, to pour and serve a drink, a robot should be able to recognize novel…

Improving the generalization capabilities of general-purpose robotic manipulation agents in the real world has long been a significant challenge. Existing approaches often rely on collecting large-scale robotic data which is costly and…

机器人学 · 计算机科学 2025-02-10 Jiange Yang , Wenhui Tan , Chuhao Jin , Keling Yao , Bei Liu , Jianlong Fu , Ruihua Song , Gangshan Wu , Limin Wang

Assistive robots operating in unstructured environments must understand not only what objects are, but what they can be used for. This requires grounding language-based action queries to objects that both afford the requested function and…

机器人学 · 计算机科学 2025-12-05 Zhou Chen , Joe Lin , Carson Bulgin , Sathyanarayanan N. Aakur

Humans excel at learning from expert demonstrations and solving their own problems. To equip intelligent robots and assistants, such as AR glasses, with this ability, it is essential to ground human hand interactions (i.e., affordances)…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Joya Chen , Difei Gao , Kevin Qinghong Lin , Mike Zheng Shou

Improving robustness of the Segment Anything Model (SAM) to input degradations is critical for its deployment in high-stakes applications such as autonomous driving and robotics. Our approach to this challenge prioritizes three key aspects:…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Sohyun Lee , Yeho Gwon , Lukas Hoyer , Suha Kwak