中文
相关论文

相关论文: WorldAfford: Affordance Grounding based on Natural…

200 篇论文

The advancement in computing power has significantly reduced the training times for deep learning, fostering the rapid development of networks designed for object recognition. However, the exploration of object utility, which is the…

计算机视觉与模式识别 · 计算机科学 2024-07-23 İsmail Özçil , A. Buğra Koku

Despite significant advancements in text-to-motion synthesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Zan Wang , Yixin Chen , Baoxiong Jia , Puhao Li , Jinlu Zhang , Jingze Zhang , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Jiajin Tang , Zhengxuan Wei , Ge Zheng , Sibei Yang

Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Lian He , Meng Liu , Qilang Ye , Yu Zhou , Xiang Deng , Gangyi Ding

Affordance detection refers to identifying the potential action possibilities of objects in an image, which is a crucial ability for robot perception and manipulation. To empower robots with this ability in unseen scenarios, we first study…

计算机视觉与模式识别 · 计算机科学 2021-08-10 Wei Zhai , Hongchen Luo , Jing Zhang , Yang Cao , Dacheng Tao

Perceiving potential ``action possibilities'' (\ie, affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions.…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Hongchen Luo , Wei Zhai , Jiao Wang , Yang Cao , Zheng-Jun Zha

Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recognition. Existing methods, mostly following the paradigm of 3D…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Yu Huang , Zelin Peng , Changsong Wen , Xiaokang Yang , Wei Shen

Currently, task-oriented grasp detection approaches are mostly based on pixel-level affordance detection and semantic segmentation. These pixel-level approaches heavily rely on the accuracy of a 2D affordance mask, and the generated grasp…

机器人学 · 计算机科学 2022-10-18 Wenkai Chen , Hongzhuo Liang , Zhaopeng Chen , Fuchun Sun , Jianwei Zhang

Mapping and localization are two essential tasks for mobile robots in real-world applications. However, largescale and dynamic scenes challenge the accuracy and robustness of most current mature solutions. This situation becomes even worse…

机器人学 · 计算机科学 2022-01-19 Fan Wang , Chaofan Zhang , Fulin Tang , Hongkui Jiang , Yihong Wu , Yong Liu

Controlling embodied agents with many actuated degrees of freedom is a challenging task. We propose a method that can discover and interpolate between context dependent high-level actions or body-affordances. These provide an abstract,…

人工智能 · 计算机科学 2017-08-16 Nicholas Guttenberg , Martin Biehl , Ryota Kanai

In model-based learning, the agent learns behaviors by simulating trajectories based on world model predictions. Standard world models typically learn a stationary transition function that maps states and actions to next states, when an…

人工智能 · 计算机科学 2026-05-11 Qinshi Zhang , Weipeng Deng , Zhihan Jiang , Jiaming Qu , Qianren Li , Weitao Xu , Ray LC

Affordances, a foundational concept in human-computer interaction and design, have traditionally been explained by direct-perception theories, which assume that individuals perceive action possibilities directly from the environment.…

人机交互 · 计算机科学 2025-01-22 Yi-Chi Liao , Christian Holz

We present SAGA, a versatile and adaptive framework for visuomotor control that can generalize across various environments, task objectives, and user specifications. To efficiently learn such capability, our key idea is to disentangle…

机器人学 · 计算机科学 2025-12-16 Kuan Fang , Yuxin Chen , Xinghao Zhu , Farzad Niroui , Lingfeng Sun , Jiuguang Wang

The concept of affordance is important to understand the relevance of object parts for a certain functional interaction. Affordance types generalize across object categories and are not mutually exclusive. This makes the segmentation of…

计算机视觉与模式识别 · 计算机科学 2017-07-11 Johann Sawatzky , Juergen Gall

Grounding textual phrases in visual content is a meaningful yet challenging problem with various potential applications such as image-text inference or text-driven multimedia interaction. Most of the current existing methods adopt the…

计算机视觉与模式识别 · 计算机科学 2018-05-03 Zhiyuan Fang , Shu Kong , Tianshu Yu , Yezhou Yang

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances,…

机器人学 · 计算机科学 2026-01-06 Tzu-Jung Lin , Jia-Fong Yeh , Hung-Ting Su , Chung-Yi Lin , Yi-Ting Chen , Winston H. Hsu

Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We propose an approach to learn human-object interaction…

计算机视觉与模式识别 · 计算机科学 2019-06-06 Tushar Nagarajan , Christoph Feichtenhofer , Kristen Grauman

The term "affordance" denotes the behavioral meaning of objects. We propose a cognitive architecture for the detection of affordances in the visual modality. This model is based on the internal simulation of movement sequences. For each…

人工智能 · 计算机科学 2016-11-02 Wolfram Schenck , Hendrik Hasenbein , Ralf Möller

Robots need to understand their environment to perform their task. If it is possible to pre-program a visual scene analysis process in closed environments, robots operating in an open environment would benefit from the ability to learn it…

机器人学 · 计算机科学 2019-03-12 Leni K. Le Goff , Oussama Yaakoubi , Alexandre Coninx , Stephane Doncieux

Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong performance in vision-language tasks, their ability to…

计算与语言 · 计算机科学 2025-08-05 Junying Wang , Wenzhe Li , Yalun Wu , Yingji Liang , Yijin Guo , Chunyi Li , Haodong Duan , Zicheng Zhang , Guangtao Zhai