中文
相关论文

相关论文: SAGA: Open-World Mobile Manipulation via Structure…

200 篇论文

When humans perform a task with an articulated object, they interact with the object only in a handful of ways, while the space of all possible interactions is nearly endless. This is because humans have prior knowledge about what…

计算机视觉与模式识别 · 计算机科学 2023-05-30 Liquan Wang , Nikita Dvornik , Rafael Dubeau , Mayank Mittal , Animesh Garg

The reliance on language in Vision-Language-Action (VLA) models introduces ambiguity, cognitive overhead, and difficulties in precise object identification and sequential task execution, particularly in environments with multiple visually…

机器人学 · 计算机科学 2026-03-02 Donggeon Kim , Seungwon Jan , Hyeonjun Park , Daegyu Lim

Recent advances in vision, language, and multimodal learning have substantially accelerated progress in robotic foundation models, with robot manipulation remaining a central and challenging problem. This survey examines robot manipulation…

A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Jiaming Liu , Mengzhen Liu , Zhenyu Wang , Pengju An , Xiaoqi Li , Kaichen Zhou , Senqiao Yang , Renrui Zhang , Yandong Guo , Shanghang Zhang

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability…

计算机视觉与模式识别 · 计算机科学 2022-03-21 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

Semantic grasping is the problem of selecting stable grasps that are functionally suitable for specific object manipulation tasks. In order for robots to effectively perform object manipulation, a broad sense of contexts, including object…

机器人学 · 计算机科学 2020-06-09 Weiyu Liu , Angel Daruna , Sonia Chernova

Robots operating in human-centered environments should have the ability to understand how objects function: what can be done with each object, where this interaction may occur, and how the object is used to achieve a goal. To this end, we…

机器人学 · 计算机科学 2022-03-02 Jessica Borja-Diaz , Oier Mees , Gabriel Kalweit , Lukas Hermann , Joschka Boedecker , Wolfram Burgard

Robotic non-destructive disassembly of mating parts remains challenging due to the need for flexible manipulation and the limited visibility of internal structures. This study presents an affordance-guided teleoperation system that enables…

机器人学 · 计算机科学 2025-08-11 Gen Sako , Takuya Kiyokawa , Kensuke Harada , Tomoki Ishikura , Naoya Miyaji , Genichiro Matsuda

Humans seamlessly fuse anticipatory planning with immediate feedback to perform successive mobile manipulation tasks without stopping, achieving both high efficiency and reliability. Replicating this fluid and reliable behavior in robots…

机器人学 · 计算机科学 2026-04-08 Chengkai Wu , Ruilin Wang , Yixin Zeng , Jiayuan Wang , Mingjie Zhang , Guiyong Zheng , Qun Niu , Juepeng Zheng , Jun Ma , Boyu Zhou

Indoor mobile manipulation (MoMA) enables robots to translate natural language instructions into physical actions, yet long-horizon execution remains challenging due to cascading errors and limited generalization across diverse…

机器人学 · 计算机科学 2026-03-10 Chenxu Li , Zixuan Chen , Yetao Li , Jiapeng Xu , Hongyu Ding , Jieqi Shi , Jing Huo , Yang Gao

Learning to manipulate dynamic and deformable objects from a single demonstration video holds great promise in terms of scalability. Previous approaches have predominantly focused on either replaying object relationships or actor…

机器人学 · 计算机科学 2024-09-24 Jianren Wang , Kangni Liu , Dingkun Guo , Xian Zhou , Christopher G Atkeson

Interactions with articulated objects are a challenging but important task for mobile robots. To tackle this challenge, we propose a novel closed-loop control pipeline, which integrates manipulation priors from affordance estimation with…

机器人学 · 计算机科学 2023-02-07 Giulio Schiavi , Paula Wulkop , Giuseppe Rizzi , Lionel Ott , Roland Siegwart , Jen Jen Chung

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of actionable affordances. However, transferring such knowledge…

机器人学 · 计算机科学 2025-05-20 Teli Ma , Jia Zheng , Zifan Wang , Ziyao Gao , Jiaming Zhou , Junwei Liang

Traditional autonomous vehicle pipelines that follow a modular approach have been very successful in the past both in academia and industry, which has led to autonomy deployed on road. Though this approach provides ease of interpretation,…

机器学习 · 计算机科学 2021-01-18 Tanmay Agarwal , Hitesh Arora , Jeff Schneider

Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occluded, reflective, and visually ambiguous. Recent systems therefore combine multiple skills…

机器人学 · 计算机科学 2026-05-11 Haojian Huang , Jiahao Shi , Yinchuan Li , Yingcong Chen

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in robotic manipulation,enabling robots to execute natural language commands through end-to-end learning from visual observations.However, deploying large-scale…

机器人学 · 计算机科学 2025-12-16 Abdullah Yahya Abdullah Omaisan , Ibrahim Sheikh Mohamed

Deploying robots in open-ended unstructured environments such as homes has been a long-standing research problem. However, robots are often studied only in closed-off lab settings, and prior mobile manipulation work is restricted to…

机器人学 · 计算机科学 2024-01-30 Haoyu Xiong , Russell Mendonca , Kenneth Shaw , Deepak Pathak

Building general-purpose robots to perform a diverse range of tasks in a large variety of environments in the physical world at the human level is extremely challenging. It requires the robot learning to be sample-efficient, generalizable,…

机器人学 · 计算机科学 2022-03-03 Jun Lv , Qiaojun Yu , Lin Shao , Wenhai Liu , Wenqiang Xu , Cewu Lu

We present MUG, a novel interactive task for multimodal grounding where a user and an agent work collaboratively on an interface screen. Prior works modeled multimodal UI grounding in one round: the user gives a command and the agent…

计算与语言 · 计算机科学 2022-10-03 Tao Li , Gang Li , Jingjie Zheng , Purple Wang , Yang Li

Robot learning approaches such as behavior cloning and reinforcement learning have shown great promise in synthesizing robot skills from human demonstrations in specific environments. However, these approaches often require task-specific…

机器人学 · 计算机科学 2025-04-09 Arthur Bucker , Pablo Ortega-Kral , Jonathan Francis , Jean Oh