English
Related papers

Related papers: Afford-X: Generalizable and Slim Affordance Reason…

200 papers

To be capable of lifelong learning in a real-life environment, robots have to tackle multiple challenges. Being able to relate physical properties they may observe in their environment to possible interactions they may have is one of them.…

Artificial Intelligence · Computer Science 2020-09-24 Alexandre Manoury , Sao Mai Nguyen , Cédric Buche

Affordances are the potential actions an agent can perform on an object, as observed by a camera. Visual affordance prediction is formulated differently for tasks such as grasping detection, affordance classification, affordance…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Tommaso Apicella , Alessio Xompero , Andrea Cavallaro

Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructions into robotic actions. Despite their recent advancements,…

Robotics · Computer Science 2025-05-27 Tuan Van Vo , Tan Quang Nguyen , Khang Minh Nguyen , Duy Ho Minh Nguyen , Minh Nhat Vu

We present SAGA, a versatile and adaptive framework for visuomotor control that can generalize across various environments, task objectives, and user specifications. To efficiently learn such capability, our key idea is to disentangle…

Robotics · Computer Science 2025-12-16 Kuan Fang , Yuxin Chen , Xinghao Zhu , Farzad Niroui , Lingfeng Sun , Jiuguang Wang

A generalist robot equipped with learned skills must be able to perform many tasks in many different environments. However, zero-shot generalization to new settings is not always possible. When the robot encounters a new environment or…

Robotics · Computer Science 2021-06-15 Alexander Khazatsky , Ashvin Nair , Daniel Jing , Sergey Levine

Open-Vocabulary 3D object affordance grounding aims to anticipate ``action possibilities'' regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yawen Shao , Wei Zhai , Yuhang Yang , Hongchen Luo , Yang Cao , Zheng-Jun Zha

Reinforcement learning (RL) is a promising approach for robotic manipulation, but it can suffer from low sample efficiency and requires extensive exploration of large state-action spaces. Recent methods leverage the commonsense knowledge…

Robotics · Computer Science 2026-04-15 Jelle Luijkx , Runyu Ma , Zlatan Ajanović , Jens Kober

The physical and textural attributes of objects have been widely studied for recognition, detection and segmentation tasks in computer vision.~A number of datasets, such as large scale ImageNet, have been proposed for feature learning using…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Zeyad Khalifa , Syed Afaq Ali Shah

Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant…

Computation and Language · Computer Science 2025-10-14 Shiqi Chen , Tongyao Zhu , Ruochen Zhou , Jinghan Zhang , Siyang Gao , Juan Carlos Niebles , Mor Geva , Junxian He , Jiajun Wu , Manling Li

Enabling robots to explore and act in unfamiliar environments under ambiguous human instructions by interactively identifying task-relevant objects (e.g., identifying cups or beverages for "I'm thirsty") remains challenging for existing…

Robotics · Computer Science 2026-02-06 Hengxuan Xu , Fengbo Lan , Zhixin Zhao , Shengjie Wang , Mengqiao Liu , Jieqian Sun , Yu Cheng , Tao Zhang

In this study, we address the problem of open-vocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on free-form natural language instructions. This task is challenging, as it…

Robotics · Computer Science 2025-12-23 Ryosuke Korekata , Quanting Xie , Yonatan Bisk , Komei Sugiura

Localizing functional regions of objects or affordances is an important aspect of scene understanding. In this work, we cast the problem of affordance segmentation as that of semantic image segmentation. In order to explore various levels…

Computer Vision and Pattern Recognition · Computer Science 2016-08-01 Abhilash Srikantha , Juergen Gall

With the rapid progress of Large Language Models (LLMs), it becomes increasingly important to understand their abilities and limitations. In two experiments, we investigate the causal and compositional reasoning abilities of LLMs and humans…

Computation and Language · Computer Science 2025-02-27 Magnus F. Gjerde , Vanessa Cheung , David Lagnado

Achieving precise and generalizable grasping across diverse objects and environments is essential for intelligent and collaborative robotic systems. However, existing approaches often struggle with ambiguous affordance reasoning and limited…

Robotics · Computer Science 2025-03-11 Ruixiang Wang , Huayi Zhou , Xinyue Yao , Guiliang Liu , Kui Jia

Humans excel at acquiring knowledge through observation. For example, we can learn to use new tools by watching demonstrations. This skill is fundamental for intelligent systems to interact with the world. A key step to acquire this skill…

Computer Vision and Pattern Recognition · Computer Science 2023-03-20 Gen Li , Varun Jampani , Deqing Sun , Laura Sevilla-Lara

An autonomous robot should be able to evaluate the affordances that are offered by a given situation. Here we address this problem by designing a system that can densely predict affordances given only a single 2D RGB image. This is achieved…

Computer Vision and Pattern Recognition · Computer Science 2017-09-27 Timo Lüddecke , Florentin Wörgötter

In order to *generalize* to various tasks in the wild, robotic agents will need a suitable representation (i.e., vision network) that enables the robot to predict optimal actions given high dimensional vision inputs. However, learning such…

Robotics · Computer Science 2024-07-29 Mohan Kumar Srirama , Sudeep Dasari , Shikhar Bahl , Abhinav Gupta

Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance alignment efficiency, primarily through minimal-data…

Computation and Language · Computer Science 2025-06-19 Hao Chen , Haoze Li , Zhiqing Xiao , Lirong Gao , Qi Zhang , Xiaomeng Hu , Ningtao Wang , Xing Fu , Junbo Zhao

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn…

Interactive articulated manipulation requires long-horizon, multi-step interactions with appliances while maintaining physical consistency. Existing vision-language and diffusion-based policies struggle to generalize across parts,…

Robotics · Computer Science 2025-12-01 Yuhan Wu , Tiantian Wei , Shuo Wang , ZhiChao Wang , Yanyong Zhang , Daniel Cremers , Yan Xia