中文
相关论文

相关论文: RAIL: Robot Affordance Imagination with Large Lang…

200 篇论文

Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance grounding teaches agents…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Ji Ha Jang , Hoigi Seo , Se Young Chun

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding human-object interactions, but their application to robotic systems with non-humanoid morphologies remains largely unexplored. This work investigates…

机器人学 · 计算机科学 2026-04-22 Jess Jones , Raul Santos-Rodriguez , Sabine Hauert

Avoiding obstacles in the perceived world has been the classical approach to autonomous mobile robot navigation. However, this usually leads to unnatural and inefficient motions that significantly differ from the way humans move in tight…

机器人学 · 计算机科学 2021-02-10 Maozhen Wang , Rui Luo , Aykut Ozgun Onol , Taskin Padir

In order to *generalize* to various tasks in the wild, robotic agents will need a suitable representation (i.e., vision network) that enables the robot to predict optimal actions given high dimensional vision inputs. However, learning such…

机器人学 · 计算机科学 2024-07-29 Mohan Kumar Srirama , Sudeep Dasari , Shikhar Bahl , Abhinav Gupta

Rigged objects are commonly used in artist pipelines, as they can flexibly adapt to different scenes and postures. However, articulating the rigs into realistic affordance-aware postures (e.g., following the context, respecting the physics…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Yu-Chu Yu , Chieh Hubert Lin , Hsin-Ying Lee , Chaoyang Wang , Yu-Chiang Frank Wang , Ming-Hsuan Yang

When told to "cut the cake," a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-world scenes, multiple objects may share identical affordances, yet only one is appropriate…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jingliang Li , Jindou Jia , Tuo An , Chuhao Zhou , Xiangyu Chen , Shilin Shan , Boyu Ma , Bofan Lyu , Gen Li , Jianfei Yang

In this paper, we present an approach for robot learning of social affordance from human activity videos. We consider the problem in the context of human-robot interaction: Our approach learns structural representations of human-human (and…

机器人学 · 计算机科学 2016-04-22 Tianmin Shu , M. S. Ryoo , Song-Chun Zhu

Perceiving potential ``action possibilities'' (\ie, affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions.…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Hongchen Luo , Wei Zhai , Jiao Wang , Yang Cao , Zheng-Jun Zha

Today, there are two major paradigms for vision-based autonomous driving systems: mediated perception approaches that parse an entire scene to make a driving decision, and behavior reflex approaches that directly map an input image to a…

计算机视觉与模式识别 · 计算机科学 2015-09-29 Chenyi Chen , Ari Seff , Alain Kornhauser , Jianxiong Xiao

In this study, we explore the sophisticated domain of task planning for robust household embodied agents, with a particular emphasis on the intricate task of selecting substitute objects. We introduce the CommonSense Object Affordance Task…

人工智能 · 计算机科学 2024-10-24 Ayush Agrawal , Raghav Prabhakar , Anirudh Goyal , Dianbo Liu

Robotic manipulation faces critical challenges in understanding spatial affordances--the "where" and "how" of object interactions--essential for complex manipulation tasks like wiping a board or stacking objects. Existing methods, including…

Perceiving and manipulating 3D articulated objects in diverse environments is essential for home-assistant robots. Recent studies have shown that point-level affordance provides actionable priors for downstream manipulation tasks. However,…

机器人学 · 计算机科学 2025-09-17 Ruihai Wu , Kai Cheng , Yan Shen , Chuanruo Ning , Guanqi Zhan , Hao Dong

Garment manipulation has attracted increasing attention due to its critical role in home-assistant robotics. However, the majority of existing garment manipulation works assume an initial state consisting of only one garment, while piled…

机器人学 · 计算机科学 2026-03-05 Mingleyang Li , Yuran Wang , Yue Chen , Tianxing Chen , Jiaqi Liang , Zishun Shen , Haoran Lu , Ruihai Wu , Hao Dong

Affordances - i.e. possibilities for action that an environment or objects in it provide - are important for robots operating in human environments to perceive. Existing approaches train such capabilities on annotated static images or…

机器人学 · 计算机科学 2025-07-24 Rachel Ringe , Mihai Pomarlan , Nikolaos Tsiogkas , Stefano De Giorgis , Maria Hedblom , Rainer Malaka

Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded solutions in open-ended environments, beyond pattern…

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Dennis Rotondi , Fabio Scaparro , Hermann Blum , Kai O. Arras

Grasping occluded objects in cluttered environments is an essential component in complex robotic manipulation tasks. In this paper, we introduce an AffordanCE-driven Next-Best-View planning policy (ACE-NBV) that tries to find a feasible…

In this work, we focus on the task of weakly supervised affordance grounding, where a model is trained to identify affordance regions on objects using human-object interaction images and egocentric object images without dense labels.…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Peiran Xu , Yadong Mu

Constructing a diverse repertoire of manipulation skills in a scalable fashion remains an unsolved challenge in robotics. One way to address this challenge is with unstructured human play, where humans operate freely in an environment to…

机器人学 · 计算机科学 2022-10-24 Suneel Belkhale , Dorsa Sadigh

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Yufei Ye , Xueting Li , Abhinav Gupta , Shalini De Mello , Stan Birchfield , Jiaming Song , Shubham Tulsiani , Sifei Liu