中文
相关论文

相关论文: AffordanceLLM: Grounding Affordance from Vision La…

200 篇论文

Enabling humans and robots to collaborate effectively requires purposeful communication and an understanding of each other's affordances. Prior work in human-robot collaboration has incorporated knowledge of human affordances, i.e., their…

机器人学 · 计算机科学 2023-12-22 Drake Moore , Mark Zolotas , Taskin Padir

Humans show an innate capability to identify tools to support specific actions. The association between objects parts and the actions they facilitate is usually named affordance. Being able to segment objects parts depending on the tasks…

计算机视觉与模式识别 · 计算机科学 2024-04-19 Claudia Cuttano , Gabriele Rosi , Gabriele Trivigno , Giuseppe Averta

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances,…

机器人学 · 计算机科学 2026-01-06 Tzu-Jung Lin , Jia-Fong Yeh , Hung-Ting Su , Chung-Yi Lin , Yi-Ting Chen , Winston H. Hsu

Affordance, defined as the potential actions that an object offers, is crucial for embodied AI agents. For example, such knowledge directs an agent to grasp a knife by the handle for cutting or by the blade for safe handover. While existing…

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Dongming Wu , Yanping Fu , Saike Huang , Yingfei Liu , Fan Jia , Nian Liu , Feng Dai , Tiancai Wang , Rao Muhammad Anwer , Fahad Shahbaz Khan , Jianbing Shen

Learning to understand and infer object functionalities is an important step towards robust visual intelligence. Significant research efforts have recently focused on segmenting the object parts that enable specific types of human-object…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Spyridon Thermos , Petros Daras , Gerasimos Potamianos

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determining where to interact in complex visual scenes. While…

机器人学 · 计算机科学 2026-05-26 Runze Wang , Yuqian Fu , Yu Li , Tao Lin , Tianwen Qian , Mohamed Elhoseiny , Bo Zhao , Yanwei Fu , Yu-Gang Jiang , Xiangyang Xue

Recent years have witnessed an emerging paradigm shift toward embodied artificial intelligence, in which an agent must learn to solve challenging tasks by interacting with its environment. There are several challenges in solving embodied…

机器人学 · 计算机科学 2022-10-26 Zhiwei Jia , Kaixiang Lin , Yizhou Zhao , Qiaozi Gao , Govind Thattai , Gaurav Sukhatme

Facilitating an entity's interaction with objects requires accurately identifying parts that afford specific actions. Weakly supervised affordance grounding (WSAG) seeks to imitate human learning from third-person demonstrations, where…

计算机视觉与模式识别 · 计算机科学 2025-08-12 WonJun Moon , Hyun Seok Seong , Jae-Pil Heo

Object affordance is an important concept in hand-object interaction, providing information on action possibilities based on human motor capacity and objects' physical property thus benefiting tasks such as action anticipation and robot…

计算机视觉与模式识别 · 计算机科学 2023-02-13 Zecheng Yu , Yifei Huang , Ryosuke Furuta , Takuma Yagi , Yusuke Goutsu , Yoichi Sato

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Dennis Rotondi , Fabio Scaparro , Hermann Blum , Kai O. Arras

Contrary to the vast literature in modeling, perceiving, and understanding agent-object (e.g., human-object, hand-object, robot-object) interaction in computer vision and robotics, very few past works have studied the task of object-object…

计算机视觉与模式识别 · 计算机科学 2021-10-27 Kaichun Mo , Yuzhe Qin , Fanbo Xiang , Hao Su , Leonidas Guibas

Reasoning about object affordances allows an autonomous agent to perform generalised manipulation tasks among object instances. While current approaches to grasp affordance estimation are effective, they are limited to a single hypothesis.…

机器人学 · 计算机科学 2019-06-25 Paola Ardón , Èric Pairet , Ronald P. A. Petrick , Subramanian Ramamoorthy , Katrin S. Lohan

Task driven object detection aims to detect object instances suitable for affording a task in an image. Its challenge lies in object categories available for the task being too diverse to be limited to a closed set of object vocabulary for…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Jiajin Tang , Ge Zheng , Jingyi Yu , Sibei Yang

Affordance refers to the perception of possible actions allowed by an object. Despite its relevance to human-computer interaction, no existing theory explains the mechanisms that underpin affordance-formation; that is, how affordances are…

人机交互 · 计算机科学 2022-01-10 Yi-Chi Liao , Kashyap Todi , Aditya Acharya , Antti Keurulainen , Andrew Howes , Antti Oulasvirta

Recent advances in task planning leverage Large Language Models (LLMs) to improve generalizability by combining such models with classical planning algorithms to address their inherent limitations in reasoning capabilities. However, these…

机器人学 · 计算机科学 2024-09-17 Timo Birr , Christoph Pohl , Abdelrahman Younes , Tamim Asfour

Many robotic tasks in real-world environments require physical interactions with an object such as pick up or push. For successful interactions, the robot needs to know the object's affordances, which are defined as the potential actions…

机器人学 · 计算机科学 2025-01-13 Paula Wulkop , Halil Umut Özdemir , Antonia Hüfner , Jen Jen Chung , Roland Siegwart , Lionel Ott

Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual representations of most VLA models are often dominated by global object appearance and struggle…

机器人学 · 计算机科学 2026-05-19 Weijie Kong , Zhian Su , Wei Yu , Huixu Dong

In this work, we address the challenge of affordance detection in 3D point clouds, a task that requires effectively capturing fine-grained alignments between point clouds and text. Existing methods often struggle to model such alignments,…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Junsei Tokumitsu , Yuiga Wada

Embodied agents operating in human spaces must be able to master how their environment works: what objects can the agent use, and how can it use them? We introduce a reinforcement learning approach for exploration for interaction, whereby…

计算机视觉与模式识别 · 计算机科学 2020-10-20 Tushar Nagarajan , Kristen Grauman