English
Related papers

Related papers: AFFORD2ACT: Affordance-Guided Automatic Keypoint S…

200 papers

Planning in realistic environments requires searching in large planning spaces. Affordances are a powerful concept to simplify this search, because they model what actions can be successful in a given situation. However, the classical…

Robotics · Computer Science 2021-06-24 Danfei Xu , Ajay Mandlekar , Roberto Martín-Martín , Yuke Zhu , Silvio Savarese , Li Fei-Fei

Autonomous manipulation in everyday tasks requires flexible action generation to handle complex, diverse real-world environments, such as objects with varying hardness and softness. Imitation Learning (IL) enables robots to learn complex…

Robotics · Computer Science 2024-12-12 Masato Kobayashi , Thanpimon Buamanee , Takumi Kobayashi

Training end-to-end deep robot policies requires a lot of domain-, task-, and hardware-specific data, which is often costly to provide. In this work, we propose to tackle this issue by employing a deep neural network with a modular…

Robotics · Computer Science 2019-03-12 Aleksi Hämäläinen , Karol Arndt , Ali Ghadirzadeh , Ville Kyrki

3D articulated objects are inherently challenging for manipulation due to the varied geometries and intricate functionalities associated with articulated objects.Point-level affordance, which predicts the per-point actionable score and thus…

Robotics · Computer Science 2024-03-08 Suhan Ling , Yian Wang , Shiguang Wu , Yuzheng Zhuang , Tianyi Xu , Yu Li , Chang Liu , Hao Dong

A robot's ability to act is fundamentally constrained by what it can perceive. Many existing approaches to visual representation learning utilize general-purpose training criteria, e.g. image reconstruction, smoothness in latent space, or…

We propose Point2Act, which directly retrieves the 3D action point relevant to a contextually described task, leveraging Multimodal Large Language Models (MLLMs). Foundation models opened the possibility for generalist robots that can…

Robotics · Computer Science 2026-03-05 Sang Min Kim , Hyeongjun Heo , Junho Kim , Yonghyeon Lee , Young Min Kim

Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are…

Affordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yizhou Huang , Fan Yang , Guoliang Zhu , Gen Li , Hao Shi , Yukun Zuo , Wenrui Chen , Zhiyong Li , Kailun Yang

Grounding object affordance is fundamental to robotic manipulation as it establishes the critical link between perception and action among interacting objects. However, prior works predominantly focus on predicting single-object affordance,…

Robotics · Computer Science 2025-09-09 Tongxuan Tian , Xuhui Kang , Yen-Ling Kuo

Imitation learning enables robots to acquire complex manipulation skills from human demonstrations, but current methods rely solely on low-level sensorimotor data while ignoring the rich semantic knowledge humans naturally possess about…

Machine Learning · Computer Science 2026-01-27 Jakob Karalus , Friedhelm Schwenker

Robotic manipulation faces critical challenges in understanding spatial affordances--the "where" and "how" of object interactions--essential for complex manipulation tasks like wiping a board or stacking objects. Existing methods, including…

Robots often struggle to follow free-form human instructions in real-world settings due to computational and sensing limitations. We address this gap with a lightweight, fully on-device pipeline that converts natural-language commands into…

Robotics · Computer Science 2026-02-11 Archit Sharma , Dharmendra Sharma , John Rebeiro , Peeyush Thakur , Narendra Dhar , Laxmidhar Behera

Robotic affordances, providing information about what actions can be taken in a given situation, can aid robotic manipulation. However, learning about affordances requires expensive large annotated datasets of interactions or…

Robotics · Computer Science 2024-06-07 Pietro Mazzaglia , Taco Cohen , Daniel Dijkman

Robotic affordances, providing information about what actions can be taken in a given situation, can aid robotic manipulation. However, learning about affordances requires expensive large annotated datasets of interactions or…

Robotics · Computer Science 2024-06-14 Pietro Mazzaglia , Taco Cohen , Daniel Dijkman

Simulation-to-decision learning enables safe policy training in digital environments without risking real-world deployment, and has become essential in mission-critical domains such as supply chains and industrial systems. However,…

Machine Learning · Computer Science 2026-03-11 Hongyu Cao , Jinghan Zhang , Kunpeng Liu , Dongjie Wang , Feng Xia , Haifeng Chen , Xiaohua Hu , Yanjie Fu

In this paper, we introduce Attention Prompt Tuning (APT) - a computationally efficient variant of prompt tuning for video-based applications such as action recognition. Prompt tuning approaches involve injecting a set of learnable prompts…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Wele Gedara Chaminda Bandara , Vishal M. Patel

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance generalization by improving the visual feature encoding for…

Robotics · Computer Science 2025-03-21 Shijie Wu , Yihang Zhu , Yunao Huang , Kaizhen Zhu , Jiayuan Gu , Jingyi Yu , Ye Shi , Jingya Wang

Enabling robots to explore and act in unfamiliar environments under ambiguous human instructions by interactively identifying task-relevant objects (e.g., identifying cups or beverages for "I'm thirsty") remains challenging for existing…

Robotics · Computer Science 2026-02-06 Hengxuan Xu , Fengbo Lan , Zhixin Zhao , Shengjie Wang , Mengqiao Liu , Jieqian Sun , Yu Cheng , Tao Zhang

Visual affordance learning is crucial for robots to understand and interact effectively with the physical world. Recent advances in this field attempt to leverage pre-trained knowledge of vision-language foundation models to learn…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Qian Zhang , Lin Zhang , Xing Fang , Mingxin Zhang , Zhiyuan Wei , Ran Song , Wei Zhang

Robots operating in human-centered environments should have the ability to understand how objects function: what can be done with each object, where this interaction may occur, and how the object is used to achieve a goal. To this end, we…