English
Related papers

Related papers: Learning Visual Affordance Grounding from Demonstr…

200 papers

We investigate a new problem of detecting hands and recognizing their physical contact state in unconstrained conditions. This is a challenging inference task given the need to reason beyond the local appearance of hands. The lack of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Supreeth Narasimhaswamy , Trung Nguyen , Minh Hoai

Since Convolutional Neural Networks (ConvNets) are able to simultaneously learn features and classifiers to discriminate different categories of activities, recent works have employed ConvNets approaches to perform human activity…

Computer Vision and Pattern Recognition · Computer Science 2018-11-19 Artur Jordao , Ricardo Kloss , William Robson Schwartz

Robotic grasping of house-hold objects has made remarkable progress in recent years. Yet, human grasps are still difficult to synthesize realistically. There are several key reasons: (1) the human hand has many degrees of freedom (more than…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 Korrawe Karunratanakul , Jinlong Yang , Yan Zhang , Michael Black , Krikamol Muandet , Siyu Tang

Many vision and language models suffer from poor visual grounding - often falling back on easy-to-learn language priors rather than basing their decisions on visual concepts in the image. In this work, we propose a generic approach called…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ramprasaath R. Selvaraju , Stefan Lee , Yilin Shen , Hongxia Jin , Shalini Ghosh , Larry Heck , Dhruv Batra , Devi Parikh

We introduce AffordanceGrasp-R1, a reasoning-driven affordance segmentation framework for robotic grasping that combines a chain-of-thought (CoT) cold-start strategy with reinforcement learning to enhance deduction and spatial grounding. In…

Robotics · Computer Science 2026-02-04 Dingyi Zhou , Mu He , Zhuowei Fang , Xiangtong Yao , Yinlong Liu , Alois Knoll , Hu Cao

We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algorithms to address VOG, fail to explore the object relation…

Computer Vision and Pattern Recognition · Computer Science 2020-03-25 Arka Sadhu , Kan Chen , Ram Nevatia

A significant challenge for real-world robotic manipulation is the effective 6DoF grasping of objects in cluttered scenes from any single viewpoint without the need for additional scene exploration. This work reinterprets grasping as…

Robotics · Computer Science 2024-05-30 Snehal Jauhri , Ishikaa Lunawat , Georgia Chalvatzaki

This paper tackles the problem of semi-supervised video object segmentation, that is, segmenting an object in a sequence given its mask in the first frame. One of the main challenges in this scenario is the change of appearance of the…

Computer Vision and Pattern Recognition · Computer Science 2018-07-19 Sergi Caelles , Yuhua Chen , Jordi Pont-Tuset , Luc Van Gool

Affordances represent the inherent effect and action possibilities that objects offer to the agents within a given context. From a theoretical viewpoint, affordances bridge the gap between effect and action, providing a functional…

Robotics · Computer Science 2024-10-11 Hakan Aktas , Yukie Nagai , Minoru Asada , Matteo Saveriano , Erhan Oztop , Emre Ugur

Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pairs. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Zhu Zhang , Zhou Zhao , Zhijie Lin , Baoxing Huai , Nicholas Jing Yuan

Video object segmentation is an essential task in robot manipulation to facilitate grasping and learning affordances. Incremental learning is important for robotics in unstructured environments, since the total number of objects and their…

Computer Vision and Pattern Recognition · Computer Science 2019-03-14 Mennatullah Siam , Chen Jiang , Steven Lu , Laura Petrich , Mahmoud Gamal , Mohamed Elhoseiny , Martin Jagersand

Recent graph convolutional neural networks (GCNs) have shown high performance in the field of human action recognition by using human skeleton poses. However, it fails to detect human-object interaction cases successfully due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Hesham M. Shehata , Mohammad Abdolrahmani

Natural Human-Robot Interaction (HRI) is one of the key components for service robots to be able to work in human-centric environments. In such dynamic environments, the robot needs to understand the intention of the user to accomplish a…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Giorgos Tziafas , Hamidreza Kasaei

Recent advances in saliency detection have utilized deep learning to obtain high level features to detect salient regions in a scene. These advances have demonstrated superior results over previous works that utilize hand-crafted low level…

Computer Vision and Pattern Recognition · Computer Science 2016-04-20 Gayoung Lee , Yu-Wing Tai , Junmo Kim

In recent years, there has been a renewed interest in jointly modeling perception and action. At the core of this investigation is the idea of modeling affordances(Affordances are opportunities of interaction in the scene. In other words,…

Computer Vision and Pattern Recognition · Computer Science 2018-04-10 Xiaolong Wang , Rohit Girdhar , Abhinav Gupta

Robots operating in human-centered environments should have the ability to understand how objects function: what can be done with each object, where this interaction may occur, and how the object is used to achieve a goal. To this end, we…

What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Qing Zhang , Xuesong Li , Jing Zhang

Cloth in the real world is often crumpled, self-occluded, or folded in on itself such that key regions, such as corners, are not directly graspable, making manipulation difficult. We propose a system that leverages visual and tactile…

Robotics · Computer Science 2022-12-13 Neha Sunil , Shaoxiong Wang , Yu She , Edward Adelson , Alberto Rodriguez

Visual Grounding (VG) aims to locate the most relevant region in an image, based on a flexible natural language query but not a pre-defined label, thus it can be a more useful technique than object detection in practice. Most…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Chaorui Deng , Qi Wu , Guanghui Xu , Zhuliang Yu , Yanwu Xu , Kui Jia , Mingkui Tan

Short Term object-interaction Anticipation consists in detecting the location of the next active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric video. This ability…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Lorenzo Mur Labadia , Ruben Martinez-Cantin , Jose J. Guerrero , Giovanni M. Farinella , Antonino Furnari