English
Related papers

Related papers: Learning Visual Affordance Grounding from Demonstr…

200 papers

The video grounding (VG) task aims to locate the queried action or event in an untrimmed video based on rich linguistic descriptions. Existing proposal-free methods are trapped in complex interaction between video and query, overemphasizing…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Kun Li , Dan Guo , Meng Wang

This paper presents HoughNet, a one-stage, anchor-free, voting-based, bottom-up object detection method. Inspired by the Generalized Hough Transform, HoughNet determines the presence of an object at a certain location by the sum of the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Nermin Samet , Samet Hicsonmez , Emre Akbas

Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision…

Robotics · Computer Science 2026-04-14 Jiawei Zhang , Kaizhe Hu , Yingqian Huang , Yuanchen Ju , Zhengrong Xue , Huazhe Xu

We explore how intermediate policy representations can facilitate generalization by providing guidance on how to perform manipulation tasks. Existing representations such as language, goal images, and trajectory sketches have been shown to…

Robotics · Computer Science 2024-11-06 Soroush Nasiriany , Sean Kirmani , Tianli Ding , Laura Smith , Yuke Zhu , Danny Driess , Dorsa Sadigh , Ted Xiao

In this abstract we describe recent [4,7] and latest work on the determination of affordances in visually perceived 3D scenes. Our method builds on the hypothesis that geometry on its own provides enough information to enable the detection…

Computer Vision and Pattern Recognition · Computer Science 2019-06-14 Eduardo Ruiz , Walterio Mayol-Cuevas

Robot-assisted feeding requires reliable bite acquisition, a challenging task due to the complex interactions between utensils and food with diverse physical properties. These interactions are further complicated by the temporal variability…

Robotics · Computer Science 2025-09-03 Zhanxin Wu , Bo Ai , Tom Silver , Tapomayukh Bhattacharjee

This paper presents HoughNet, a one-stage, anchor-free, voting-based, bottom-up object detection method. Inspired by the Generalized Hough Transform, HoughNet determines the presence of an object at a certain location by the sum of the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Nermin Samet , Samet Hicsonmez , Emre Akbas

3D affordance reasoning, the task of associating human instructions with the functional regions of 3D objects, is a critical capability for embodied agents. Current methods based on 3D Gaussian Splatting (3DGS) are fundamentally limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Yuhui Zheng , Mingtao Feng , Guangming Shi

Short-Term object-interaction Anticipation consists of detecting the location of the next-active objects, the noun and verb categories of the interaction, and the time to contact from the observation of egocentric video. This ability is…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Lorenzo Mur-Labadia , Ruben Martinez-Cantin , Josechu Guerrero , Giovanni Maria Farinella , Antonino Furnari

Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges…

Computer Vision and Pattern Recognition · Computer Science 2018-09-21 Fabien Baradel , Natalia Neverova , Christian Wolf , Julien Mille , Greg Mori

How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambiguities by leveraging contextual knowledge -- such as affordances, where an object's shape and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Naru Suzuki , Takehiko Ohkawa , Tatsuro Banno , Jihyun Lee , Ryosuke Furuta , Yoichi Sato

Multi-exposure High Dynamic Range (HDR) imaging is a challenging task when facing truncated texture and complex motion. Existing deep learning-based methods have achieved great success by either following the alignment and fusion pipeline…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Lingtong Kong , Bo Li , Yike Xiong , Hao Zhang , Hong Gu , Jinwei Chen

Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and recognize their interactions. While recent two-stage approaches have made significant progress, they still face challenges due to incomplete…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Zhehao Li , Yucheng Qian , Chong Wang , Yinghao Lu , Zhihao Yang , Jiafei Wu

This paper presents an approach for learning invariant features for object affordance understanding. One of the major problems for a robotic agent acquiring a deeper understanding of affordances is finding sensory-grounded semantics. Being…

Robotics · Computer Science 2019-01-31 Martin Hjelm , Carl Henrik Ek , Renaud Detry , Danica Kragic

Can a video generation model be repurposed as an interactive world simulator? We explore the affordance perception potential of text-to-video models by teaching them to predict human-environment interaction. Given a scene image and a prompt…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Mengyi Shan , Zecheng He , Haoyu Ma , Felix Juefei-Xu , Peizhao Zhang , Tingbo Hou , Ching-Yao Chuang

Recent salient object detection (SOD) models predominantly rely on heavyweight backbones, incurring substantial computational cost and hindering their practical application in various real-world settings, particularly on edge devices. This…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yu-Huan Wu , Wei Liu , Zi-Xuan Zhu , Zizhou Wang , Yong Liu , Liangli Zhen

Understanding how humans interact with the surrounding environment, and specifically reasoning about object interactions and affordances, is a critical challenge in computer vision, robotics, and AI. Current approaches often depend on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Harry Zhang , Luca Carlone

An autonomous robot should be able to evaluate the affordances that are offered by a given situation. Here we address this problem by designing a system that can densely predict affordances given only a single 2D RGB image. This is achieved…

Computer Vision and Pattern Recognition · Computer Science 2017-09-27 Timo Lüddecke , Florentin Wörgötter

Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands understanding videos' and queries' semantic content and the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Mattia Soldan , Mengmeng Xu , Sisi Qu , Jesper Tegner , Bernard Ghanem

This paper addresses the challenge of perceiving complete object shapes through visual perception. While prior studies have demonstrated encouraging outcomes in segmenting the visible parts of objects within a scene, amodal segmentation, in…

Robotics · Computer Science 2024-08-07 Jinyu Zhang , Yongchong Gu , Jianxiong Gao , Haitao Lin , Qiang Sun , Xinwei Sun , Xiangyang Xue , Yanwei Fu