English
Related papers

Related papers: AffordanceLLM: Grounding Affordance from Vision La…

200 papers

Imitation learning has unlocked the potential for robots to exhibit highly dexterous behaviours. However, it still struggles with long-horizon, multi-object tasks due to poor sample efficiency and limited generalisation. Existing methods…

Robotics · Computer Science 2025-09-05 Krishan Rana , Jad Abou-Chakra , Sourav Garg , Robert Lee , Ian Reid , Niko Suenderhauf

Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recognition. Existing methods, mostly following the paradigm of 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yu Huang , Zelin Peng , Changsong Wen , Xiaokang Yang , Wei Shen

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Yufei Ye , Xueting Li , Abhinav Gupta , Shalini De Mello , Stan Birchfield , Jiaming Song , Shubham Tulsiani , Sifei Liu

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as "the left most chair"…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

Garment manipulation has attracted increasing attention due to its critical role in home-assistant robotics. However, the majority of existing garment manipulation works assume an initial state consisting of only one garment, while piled…

Robotics · Computer Science 2026-03-05 Mingleyang Li , Yuran Wang , Yue Chen , Tianxing Chen , Jiaqi Liang , Zishun Shen , Haoran Lu , Ruihai Wu , Hao Dong

Affordance detection from visual input is a fundamental step in autonomous robotic manipulation. Existing solutions to the problem of affordance detection rely on convolutional neural networks. However, these networks do not consider the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Antonio Rodríguez-Sánchez , Simon Haller-Seeber , David Peer , Chris Engelhardt , Jakob Mittelberger , Matteo Saveriano

While many quality metrics exist to evaluate the quality of a grasp by itself, no clear quantification of the quality of a grasp relatively to the task the grasp is used for has been defined yet. In this paper we propose a framework to…

Robotics · Computer Science 2019-07-11 Luca Cavalli , Gianpaolo Di Pietro , Matteo Matteucci

Autonomous agents must often detect affordances: the set of behaviors enabled by a situation. Affordance detection is particularly helpful in domains with large action spaces, allowing the agent to prune its search space by avoiding futile…

Artificial Intelligence · Computer Science 2018-09-03 Nancy Fulda , Daniel Ricks , Ben Murdoch , David Wingate

How do we know that a kitchen is a kitchen by looking? Relatively little is known about how we conceptualize and categorize different visual environments. Traditional models of visual perception posit that scene categorization is achieved…

Neurons and Cognition · Quantitative Biology 2014-11-20 Michelle R. Greene , Christopher Baldassano , Andre Esteva , Diane M. Beck , Li Fei-Fei

If a robotic agent wants to exploit symbolic planning techniques to achieve some goal, it must be able to properly ground an abstract planning domain in the environment in which it operates. However, if the environment is initially unknown…

Artificial Intelligence · Computer Science 2022-04-11 Leonardo Lamanna , Luciano Serafini , Alessandro Saetti , Alfonso Gerevini , Paolo Traverso

Articulated objects pose diverse manipulation challenges for robots. Since their internal structures are not directly observable, robots must adaptively explore and refine actions to generate successful manipulation trajectories. While…

Robotics · Computer Science 2025-07-25 Xiaojie Zhang , Yuanfei Wang , Ruihai Wu , Kunqi Xu , Yu Li , Liuyu Xiang , Hao Dong , Zhaofeng He

Vision-and-Language Navigation (VLN) requires grounding instructions, such as "turn right and stop at the door", to routes in a visual environment. The actual grounding can connect language to the environment through multiple modalities,…

Computation and Language · Computer Science 2019-06-11 Ronghang Hu , Daniel Fried , Anna Rohrbach , Dan Klein , Trevor Darrell , Kate Saenko

This study investigates how text-driven object affordance, which provides prior knowledge about grasp types for each object, affects image-based grasp-type recognition in robot teaching. The researchers created labeled datasets of…

Robotics · Computer Science 2023-06-07 Naoki Wake , Daichi Saito , Kazuhiro Sasabuchi , Hideki Koike , Katsushi Ikeuchi

Humans perceive and interact with the world with the awareness of equivariance, facilitating us in manipulating different objects in diverse poses. For robotic manipulation, such equivariance also exists in many scenarios. For example, no…

Robotics · Computer Science 2024-08-08 Yue Chen , Chenrui Tie , Ruihai Wu , Hao Dong

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot…

This paper develops and evaluates a novel method that allows for the detection of affordances in a scalable and multiple-instance manner on visually recovered pointclouds. Our approach has many advantages over alternative methods, as it is…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Eduardo Ruiz , Walterio Mayol-Cuevas

3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing methods rely on a pre-defined Object Lookup Table (OLT) to query Visual Language Models (VLMs) for reasoning about object locations,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Wenyuan Huang , Zhao Wang , Zhou Wei , Ting Huang , Fang Zhao , Jian Yang , Zhenyu Zhang

Affordance modeling plays an important role in visual understanding. In this paper, we aim to predict affordances of 3D indoor scenes, specifically what human poses are afforded by a given indoor environment, such as sitting on a chair or…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Xueting Li , Sifei Liu , Kihwan Kim , Xiaolong Wang , Ming-Hsuan Yang , Jan Kautz

Deformable object manipulation in robotics presents significant challenges due to uncertainties in component properties, diverse configurations, visual interference, and ambiguous prompts. These factors complicate both perception and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Wanjun Jia , Fan Yang , Mengfei Duan , Xianchi Chen , Yinxi Wang , Yiming Jiang , Wenrui Chen , Kailun Yang , Zhiyong Li

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3D-Bench, a holistic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Tianxu Wang , Zhuofan Zhang , Ziyu Zhu , Yue Fan , Jing Xiong , Pengxiang Li , Xiaojian Ma , Qing Li