English
Related papers

Related papers: Beyond the Contact: Discovering Comprehensive Affo…

200 papers

Object affordance is an important concept in hand-object interaction, providing information on action possibilities based on human motor capacity and objects' physical property thus benefiting tasks such as action anticipation and robot…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 Zecheng Yu , Yifei Huang , Ryosuke Furuta , Takuma Yagi , Yusuke Goutsu , Yoichi Sato

What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Qing Zhang , Xuesong Li , Jing Zhang

Modeling face-to-face communication in computer vision, which focuses on recognizing and analyzing nonverbal cues and behaviors during interactions, serves as the foundation for our proposed alternative to text-based Human-AI interaction.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Dragos Costea , Alina Marcu , Cristina Lazar , Marius Leordeanu

In this paper, we present an approach for robot learning of social affordance from human activity videos. We consider the problem in the context of human-robot interaction: Our approach learns structural representations of human-human (and…

Robotics · Computer Science 2016-04-22 Tianmin Shu , M. S. Ryoo , Song-Chun Zhu

Affordance learning considers the interaction opportunities for an actor in the scene and thus has wide application in scene understanding and intelligent robotics. In this paper, we focus on contextual affordance learning, i.e., using…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Jieteng Yao , Junjie Chen , Li Niu , Bin Sheng

We propose a new dataset and a novel approach to learning hand-object interaction priors for hand and articulated object pose estimation. We first collect a dataset using visual teleoperation, where the human operator can directly play…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zehao Zhu , Jiashun Wang , Yuzhe Qin , Deqing Sun , Varun Jampani , Xiaolong Wang

Accurate affordance detection and segmentation with pixel precision is an important piece in many complex systems based on interactions, such as robots and assitive devices. We present a new approach to affordance perception which enables…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Lorenzo Mur-Labadia , Jose J. Guerrero , Ruben Martinez-Cantin

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent…

Humans excel at learning from expert demonstrations and solving their own problems. To equip intelligent robots and assistants, such as AR glasses, with this ability, it is essential to ground human hand interactions (i.e., affordances)…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Joya Chen , Difei Gao , Kevin Qinghong Lin , Mike Zheng Shou

In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Sisi Dai , Wenhao Li , Haowen Sun , Haibin Huang , Chongyang Ma , Hui Huang , Kai Xu , Ruizhen Hu

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

In order to *generalize* to various tasks in the wild, robotic agents will need a suitable representation (i.e., vision network) that enables the robot to predict optimal actions given high dimensional vision inputs. However, learning such…

Robotics · Computer Science 2024-07-29 Mohan Kumar Srirama , Sudeep Dasari , Shikhar Bahl , Abhinav Gupta

We present AROS, a one-shot learning approach that uses an explicit representation of interactions between highly-articulated human poses and 3D scenes. The approach is one-shot as the method does not require re-training to add new…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Abel Pacheco-Ortega , Walterio Mayol-Cuevas

3D articulated objects are inherently challenging for manipulation due to the varied geometries and intricate functionalities associated with articulated objects.Point-level affordance, which predicts the per-point actionable score and thus…

Robotics · Computer Science 2024-03-08 Suhan Ling , Yian Wang , Shiguang Wu , Yuzheng Zhuang , Tianyi Xu , Yu Li , Chang Liu , Hao Dong

Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-level physical…

Robotics · Computer Science 2026-05-08 Hao Wang , Shiqi Wang , Qi Liu

Grounding 3D scene affordance aims to locate interactive regions in 3D environments, which is crucial for embodied agents to interact intelligently with their surroundings. Most existing approaches achieve this by mapping semantics to 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Cuiyu Liu , Wei Zhai , Yuhang Yang , Hongchen Luo , Sen Liang , Yang Cao , Zheng-Jun Zha

Fine-grained capturing of 3D HOI boosts human activity understanding and facilitates downstream visual tasks, including action recognition, holistic scene reconstruction, and human motion synthesis. Despite its significance, existing works…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Nan Jiang , Tengyu Liu , Zhexuan Cao , Jieming Cui , Zhiyuan zhang , Yixin Chen , He Wang , Yixin Zhu , Siyuan Huang

Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from narrow Fields of View (FoV) and fragmented observations,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zixin Zhang , Chenfei Liao , Hongfei Zhang , Harold Haodong Chen , Kanghao Chen , Zichen Wen , Litao Guo , Bin Ren , Xu Zheng , Yinchuan Li , Xuming Hu , Nicu Sebe , Ying-Cong Chen

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Xintao Lv , Liang Xu , Yichao Yan , Xin Jin , Congsheng Xu , Shuwen Wu , Yifan Liu , Lincheng Li , Mengxiao Bi , Wenjun Zeng , Xiaokang Yang