English
Related papers

Related papers: SAGA: Open-World Mobile Manipulation via Structure…

200 papers

Complex manipulation tasks, such as rearrangement planning of numerous objects, are combinatorially hard problems. Existing algorithms either do not scale well or assume a great deal of prior knowledge about the environment, and few offer…

Robotics · Computer Science 2021-03-25 Vasileios Vasilopoulos , Yiannis Kantaros , George J. Pappas , Daniel E. Koditschek

Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and activities in both humans and Artificial Intelligence (AI). This capability, required for…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Xiaomeng Zhu , Yuyang Li , Leiyao Cui , Pengfei Li , Huan-ang Gao , Yixin Zhu , Hao Zhao

Language-guided robot dexterous generation enables robots to grasp and manipulate objects based on human commands. However, previous data-driven methods are hard to understand intention and execute grasping with unseen categories in the…

Robotics · Computer Science 2025-07-31 Yi-Lin Wei , Mu Lin , Yuhao Lin , Jian-Jian Jiang , Xiao-Ming Wu , Ling-An Zeng , Wei-Shi Zheng

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Dongming Wu , Yanping Fu , Saike Huang , Yingfei Liu , Fan Jia , Nian Liu , Feng Dai , Tiancai Wang , Rao Muhammad Anwer , Fahad Shahbaz Khan , Jianbing Shen

Grounding 3D object affordance seeks to locate objects' ''action possibilities'' regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Yuhang Yang , Wei Zhai , Hongchen Luo , Yang Cao , Jiebo Luo , Zheng-Jun Zha

In model-based learning, the agent learns behaviors by simulating trajectories based on world model predictions. Standard world models typically learn a stationary transition function that maps states and actions to next states, when an…

Artificial Intelligence · Computer Science 2026-05-11 Qinshi Zhang , Weipeng Deng , Zhihan Jiang , Jiaming Qu , Qianren Li , Weitao Xu , Ray LC

Robots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to…

Robotics · Computer Science 2025-03-18 Yangtao Chen , Zixuan Chen , Junhui Yin , Jing Huo , Pinzhuo Tian , Jieqi Shi , Yang Gao

Service robots are expected to autonomously and efficiently work in human-centric environments. For this type of robots, object perception and manipulation are challenging tasks due to need for accurate and real-time response. This paper…

Robotics · Computer Science 2019-04-05 S. Hamidreza Kasaei , Nima Shafii , Luis Seabra Lopes , Ana Maria Tome

Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they remain brittle under distribution shifts: when test…

Robotics · Computer Science 2025-12-09 Siyu Xu , Zijian Wang , Yunke Wang , Chenghao Xia , Tao Huang , Chang Xu

Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance grounding teaches agents…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ji Ha Jang , Hoigi Seo , Se Young Chun

Recent works have shown that Large Language Models (LLMs) can be applied to ground natural language to a wide variety of robot skills. However, in practice, learning multi-task, language-conditioned robotic skills typically requires…

Robotics · Computer Science 2023-03-09 Oier Mees , Jessica Borja-Diaz , Wolfram Burgard

Visual affordance learning is a key component for robots to understand how to interact with objects. Conventional approaches in this field rely on pre-defined objects and actions, falling short of capturing diverse interactions in realworld…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we introduce Ego3D Position Encoding…

Robotics · Computer Science 2025-05-20 Delin Qu , Haoming Song , Qizhi Chen , Yuanqi Yao , Xinyi Ye , Yan Ding , Zhigang Wang , JiaYuan Gu , Bin Zhao , Dong Wang , Xuelong Li

In this paper, we present a novel approach for learning bimanual manipulation actions from human demonstration by extracting spatial constraints between affordance regions, termed affordance constraints, of the objects involved. Affordance…

Robotics · Computer Science 2024-11-19 Björn S. Plonka , Christian Dreher , Andre Meixner , Rainer Kartmann , Tamim Asfour

This paper introduces an automatic affordance reasoning paradigm tailored to minimal semantic inputs, addressing the critical challenges of classifying and manipulating unseen classes of objects in household settings. Inspired by human…

Robotics · Computer Science 2024-06-10 Ceng Zhang , Xin Meng , Dongchen Qi , Gregory S. Chirikjian

Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are…

Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing methods of visual affordance predictions often rely on manually…

Robotics · Computer Science 2025-08-27 Yihe Tang , Wenlong Huang , Yingke Wang , Chengshu Li , Roy Yuan , Ruohan Zhang , Jiajun Wu , Li Fei-Fei

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile…

While large language models (LLMs) are successful in completing various language processing tasks, they easily fail to interact with the physical world by generating control sequences properly. We find that the main reason is that LLMs are…

Artificial Intelligence · Computer Science 2024-04-18 Guangran Cheng , Chuheng Zhang , Wenzhe Cai , Li Zhao , Changyin Sun , Jiang Bian

Artificial intelligence is essential to succeed in challenging activities that involve dynamic environments, such as object manipulation tasks in indoor scenes. Most of the state-of-the-art literature explores robotic grasping methods by…

Robotics · Computer Science 2019-05-28 Paola Ardón , Èric Pairet , Ron Petrick , Subramanian Ramamoorthy , Katrin Lohan