English
Related papers

Related papers: RAIL: Robot Affordance Imagination with Large Lang…

200 papers

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability…

Computer Vision and Pattern Recognition · Computer Science 2022-03-21 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

Flexible, goal-directed behavior is a fundamental aspect of human life. Based on the free energy minimization principle, the theory of active inference formalizes the generation of such behavior from a computational neuroscience…

Artificial Intelligence · Computer Science 2022-08-03 Fedor Scholz , Christian Gumbsch , Sebastian Otte , Martin V. Butz

Affordance understanding, the task of identifying actionable regions on 3D objects, plays a vital role in allowing robotic systems to engage with and operate within the physical world. Although Visual Language Models (VLMs) have excelled in…

Embodied agents operating in human spaces must be able to master how their environment works: what objects can the agent use, and how can it use them? We introduce a reinforcement learning approach for exploration for interaction, whereby…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Tushar Nagarajan , Kristen Grauman

Objects, in particular tools, provide several action possibilities to the agents that can act on them, which are generally associated with the term of affordances. A tool is typically designed for a specific purpose, such as driving a nail…

Robotics · Computer Science 2024-07-17 Bosong Ding , Erhan Oztop , Giacomo Spigler , Murat Kirtay

Accurate affordance detection and segmentation with pixel precision is an important piece in many complex systems based on interactions, such as robots and assitive devices. We present a new approach to affordance perception which enables…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Lorenzo Mur-Labadia , Jose J. Guerrero , Ruben Martinez-Cantin

Autonomous parking technology is a key concept within autonomous driving research. This paper will propose an imaginative autonomous parking algorithm to solve issues concerned with parking. The proposed algorithm consists of three parts:…

Robotics · Computer Science 2021-08-27 Ziyue Feng , Yu Chen , Shitao Chen , Nanning Zheng

3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordance knowledge from static cues such as language and images,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Hanqing Wang , Mingyu Liu , Xiaoyu Chen , Chengwei MA , Yiming Zhong , Wenti Yin , Yuhao Liu , Zhiqing Cui , Jiahao Yuan , Lu Dai , Zhiyuan Ma , Hui Xiong

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify…

Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing methods of visual affordance predictions often rely on manually…

Robotics · Computer Science 2025-08-27 Yihe Tang , Wenlong Huang , Yingke Wang , Chengshu Li , Roy Yuan , Ruohan Zhang , Jiajun Wu , Li Fei-Fei

We investigate the knowledge of object affordances in pre-trained language models (LMs) and pre-trained Vision-Language models (VLMs). A growing body of literature shows that PTLMs fail inconsistently and non-intuitively, demonstrating a…

Computation and Language · Computer Science 2025-09-29 Sayantan Adak , Daivik Agrawal , Animesh Mukherjee , Somak Aditya

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Yan Zhao , Ruihai Wu , Zhehuan Chen , Yourong Zhang , Qingnan Fan , Kaichun Mo , Hao Dong

In current presence or availability systems, the method of presenting a user's state often supposes an instantaneous notion of that state - for example, a visualization is rendered or an inference is made about the potential actions that…

Human-Computer Interaction · Computer Science 2007-05-23 Paul M. Aoki , Allison Woodruff

Affordance prediction, which identifies interaction regions on objects based on language instructions, is critical for embodied AI. Prevailing end-to-end models couple high-level reasoning and low-level grounding into a single monolithic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Zixin Zhang , Kanghao Chen , Hanqing Wang , Hongfei Zhang , Harold Haodong Chen , Chenfei Liao , Litao Guo , Ying-Cong Chen

Achieving precise and generalizable grasping across diverse objects and environments is essential for intelligent and collaborative robotic systems. However, existing approaches often struggle with ambiguous affordance reasoning and limited…

Robotics · Computer Science 2025-03-11 Ruixiang Wang , Huayi Zhou , Xinyue Yao , Guiliang Liu , Kui Jia

This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a text query that describes an action on the object. While existing methods predict affordance…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chunghyun Park , Seunghyeon Lee , Minsu Cho

Robotic grasping is a fundamental ability for a robot to interact with the environment. Current methods focus on how to obtain a stable and reliable grasping pose in object level, while little work has been studied on part (shape)-wise…

Robotics · Computer Science 2025-05-01 Yaoxian Song , Penglei Sun , Piaopiao Jin , Yi Ren , Yu Zheng , Zhixu Li , Xiaowen Chu , Yue Zhang , Tiefeng Li , Jason Gu

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot…

3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xianqiang Gao , Pingrui Zhang , Delin Qu , Dong Wang , Zhigang Wang , Yan Ding , Bin Zhao

We study the task of language instruction-guided robotic manipulation, in which an embodied robot is supposed to manipulate the target objects based on the language instructions. In previous studies, the predicted manipulation regions of…

Robotics · Computer Science 2024-08-27 Dayou Li , Chenkun Zhao , Shuo Yang , Lin Ma , Yibin Li , Wei Zhang