中文
相关论文

相关论文: Task-Aware Bimanual Affordance Prediction via VLM-…

200 篇论文

When interacting with objects, humans effectively reason about which regions of objects are viable for an intended action, i.e., the affordance regions of the object. They can also account for subtle differences in object regions based on…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Marvin Heidinger , Snehal Jauhri , Vignesh Prasad , Georgia Chalvatzaki

Embodied agents operating in open environments must translate high-level instructions into grounded, executable behaviors, often requiring coordinated use of both hands. While recent foundation models offer strong semantic reasoning,…

机器人学 · 计算机科学 2025-12-11 Kwang Bin Lee , Jiho Kang , Sung-Hee Lee

Many everyday objects are difficult to directly grasp (e.g., a flat iPad) or manipulate functionally (e.g., opening the cap of a pen lying on a desk). Such tasks require sequential, asymmetric coordination between two arms, where one arm…

机器人学 · 计算机科学 2026-03-24 Yan Shen , Feng Jiang , Zichen He , Xiaoqi Li , Yuchen Liu , Zhiyu Li , Ruihai Wu , Hao Dong

Mobile robot platforms will increasingly be tasked with activities that involve grasping and manipulating objects in open world environments. Affordance understanding provides a robot with means to realise its goals and execute its tasks,…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Gertjan Burghouts , Marianne Schaaphok , Michael van Bekkum , Wouter Meijer , Fieke Hillerström , Jelle van Mil

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances,…

机器人学 · 计算机科学 2026-01-06 Tzu-Jung Lin , Jia-Fong Yeh , Hung-Ting Su , Chung-Yi Lin , Yi-Ting Chen , Winston H. Hsu

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of the environment,…

机器人学 · 计算机科学 2025-11-18 Xiaoshuai Hao , Yingbo Tang , Lingfeng Zhang , Yanbiao Ma , Yunfeng Diao , Ziyu Jia , Wenbo Ding , Hangjun Ye , Long Chen

Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and activities in both humans and Artificial Intelligence (AI). This capability, required for…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Xiaomeng Zhu , Yuyang Li , Leiyao Cui , Pengfei Li , Huan-ang Gao , Yixin Zhu , Hao Zhao

Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlooked by many previous works. That is, different instructions…

机器人学 · 计算机科学 2025-08-26 Bokai Ji , Jie Gu , Xiaokang Ma , Chu Tang , Jingmin Chen , Guangxia Li

In order for robots to interact with objects effectively, they must understand the form and function of each object they encounter. Essentially, robots need to understand which actions each object affords, and where those affordances can be…

机器人学 · 计算机科学 2024-05-28 Edmond Tong , Anthony Opipari , Stanley Lewis , Zhen Zeng , Odest Chadwicke Jenkins

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding human-object interactions, but their application to robotic systems with non-humanoid morphologies remains largely unexplored. This work investigates…

机器人学 · 计算机科学 2026-04-22 Jess Jones , Raul Santos-Rodriguez , Sabine Hauert

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify…

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot…

Vision-Language Models (VLMs) have shown great success as foundational models for downstream vision and natural language applications in a variety of domains. However, these models are limited to reasoning over objects and actions currently…

机器人学 · 计算机科学 2025-06-13 Zachary Chavis , Hyun Soo Park , Stephen J. Guy

Learning to manipulate 3D objects in an interactive environment has been a challenging problem in Reinforcement Learning (RL). In particular, it is hard to train a policy that can generalize over objects with different semantic categories,…

机器人学 · 计算机科学 2022-09-28 Yiran Geng , Boshi An , Haoran Geng , Yuanpei Chen , Yaodong Yang , Hao Dong

Bimanual robotic manipulation provides significant versatility, but also presents an inherent challenge due to the complexity involved in the spatial and temporal coordination between two hands. Existing works predominantly focus on…

机器人学 · 计算机科学 2025-03-24 Kun Chu , Xufeng Zhao , Cornelius Weber , Stefan Wermter

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Yan Zhao , Ruihai Wu , Zhehuan Chen , Yourong Zhang , Qingnan Fan , Kaichun Mo , Hao Dong

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determining where to interact in complex visual scenes. While…

机器人学 · 计算机科学 2026-05-26 Runze Wang , Yuqian Fu , Yu Li , Tao Lin , Tianwen Qian , Mohamed Elhoseiny , Bo Zhao , Yanwei Fu , Yu-Gang Jiang , Xiangyang Xue

Visual affordance segmentation identifies the surfaces of an object an agent can interact with. Common challenges for the identification of affordances are the variety of the geometry and physical properties of these surfaces as well as…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Tommaso Apicella , Alessio Xompero , Edoardo Ragusa , Riccardo Berta , Andrea Cavallaro , Paolo Gastaldo

We investigate the knowledge of object affordances in pre-trained language models (LMs) and pre-trained Vision-Language models (VLMs). A growing body of literature shows that PTLMs fail inconsistently and non-intuitively, demonstrating a…

计算与语言 · 计算机科学 2025-09-29 Sayantan Adak , Daivik Agrawal , Animesh Mukherjee , Somak Aditya

Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires the comprehensive understanding of a scene in multiple aspects…

计算机视觉与模式识别 · 计算机科学 2024-04-19 Shengyi Qian , Weifeng Chen , Min Bai , Xiong Zhou , Zhuowen Tu , Li Erran Li
‹ 上一页 1 2 3 10 下一页 ›