English
Related papers

Related papers: WorldAfford: Affordance Grounding based on Natural…

200 papers

The advancement in computing power has significantly reduced the training times for deep learning, fostering the rapid development of networks designed for object recognition. However, the exploration of object utility, which is the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 İsmail Özçil , A. Buğra Koku

Despite significant advancements in text-to-motion synthesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Zan Wang , Yixin Chen , Baoxiong Jia , Puhao Li , Jinlu Zhang , Jingze Zhang , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jiajin Tang , Zhengxuan Wei , Ge Zheng , Sibei Yang

Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Lian He , Meng Liu , Qilang Ye , Yu Zhou , Xiang Deng , Gangyi Ding

Affordance detection refers to identifying the potential action possibilities of objects in an image, which is a crucial ability for robot perception and manipulation. To empower robots with this ability in unseen scenarios, we first study…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Wei Zhai , Hongchen Luo , Jing Zhang , Yang Cao , Dacheng Tao

Perceiving potential ``action possibilities'' (\ie, affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Hongchen Luo , Wei Zhai , Jiao Wang , Yang Cao , Zheng-Jun Zha

Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recognition. Existing methods, mostly following the paradigm of 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yu Huang , Zelin Peng , Changsong Wen , Xiaokang Yang , Wei Shen

Currently, task-oriented grasp detection approaches are mostly based on pixel-level affordance detection and semantic segmentation. These pixel-level approaches heavily rely on the accuracy of a 2D affordance mask, and the generated grasp…

Robotics · Computer Science 2022-10-18 Wenkai Chen , Hongzhuo Liang , Zhaopeng Chen , Fuchun Sun , Jianwei Zhang

Mapping and localization are two essential tasks for mobile robots in real-world applications. However, largescale and dynamic scenes challenge the accuracy and robustness of most current mature solutions. This situation becomes even worse…

Robotics · Computer Science 2022-01-19 Fan Wang , Chaofan Zhang , Fulin Tang , Hongkui Jiang , Yihong Wu , Yong Liu

Controlling embodied agents with many actuated degrees of freedom is a challenging task. We propose a method that can discover and interpolate between context dependent high-level actions or body-affordances. These provide an abstract,…

Artificial Intelligence · Computer Science 2017-08-16 Nicholas Guttenberg , Martin Biehl , Ryota Kanai

In model-based learning, the agent learns behaviors by simulating trajectories based on world model predictions. Standard world models typically learn a stationary transition function that maps states and actions to next states, when an…

Artificial Intelligence · Computer Science 2026-05-11 Qinshi Zhang , Weipeng Deng , Zhihan Jiang , Jiaming Qu , Qianren Li , Weitao Xu , Ray LC

Affordances, a foundational concept in human-computer interaction and design, have traditionally been explained by direct-perception theories, which assume that individuals perceive action possibilities directly from the environment.…

Human-Computer Interaction · Computer Science 2025-01-22 Yi-Chi Liao , Christian Holz

We present SAGA, a versatile and adaptive framework for visuomotor control that can generalize across various environments, task objectives, and user specifications. To efficiently learn such capability, our key idea is to disentangle…

Robotics · Computer Science 2025-12-16 Kuan Fang , Yuxin Chen , Xinghao Zhu , Farzad Niroui , Lingfeng Sun , Jiuguang Wang

The concept of affordance is important to understand the relevance of object parts for a certain functional interaction. Affordance types generalize across object categories and are not mutually exclusive. This makes the segmentation of…

Computer Vision and Pattern Recognition · Computer Science 2017-07-11 Johann Sawatzky , Juergen Gall

Grounding textual phrases in visual content is a meaningful yet challenging problem with various potential applications such as image-text inference or text-driven multimedia interaction. Most of the current existing methods adopt the…

Computer Vision and Pattern Recognition · Computer Science 2018-05-03 Zhiyuan Fang , Shu Kong , Tianshu Yu , Yezhou Yang

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances,…

Robotics · Computer Science 2026-01-06 Tzu-Jung Lin , Jia-Fong Yeh , Hung-Ting Su , Chung-Yi Lin , Yi-Ting Chen , Winston H. Hsu

Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We propose an approach to learn human-object interaction…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Tushar Nagarajan , Christoph Feichtenhofer , Kristen Grauman

The term "affordance" denotes the behavioral meaning of objects. We propose a cognitive architecture for the detection of affordances in the visual modality. This model is based on the internal simulation of movement sequences. For each…

Artificial Intelligence · Computer Science 2016-11-02 Wolfram Schenck , Hendrik Hasenbein , Ralf Möller

Robots need to understand their environment to perform their task. If it is possible to pre-program a visual scene analysis process in closed environments, robots operating in an open environment would benefit from the ability to learn it…

Robotics · Computer Science 2019-03-12 Leni K. Le Goff , Oussama Yaakoubi , Alexandre Coninx , Stephane Doncieux

Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong performance in vision-language tasks, their ability to…

Computation and Language · Computer Science 2025-08-05 Junying Wang , Wenzhe Li , Yalun Wu , Yingji Liang , Yijin Guo , Chunyi Li , Haodong Duan , Zicheng Zhang , Guangtao Zhai