中文
相关论文

相关论文: Binge Watching: Scaling Affordance Learning from S…

200 篇论文

Understanding images with people often entails understanding their \emph{interactions} with other objects or people. As such, given a novel image, a vision system ought to infer which other objects/people play an important role in a given…

计算机视觉与模式识别 · 计算机科学 2016-04-19 Chao-Yeh Chen , Kristen Grauman

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Chunlin Yu , Hanqing Wang , Ye Shi , Haoyang Luo , Sibei Yang , Jingyi Yu , Jingya Wang

In this study, we explore the sophisticated domain of task planning for robust household embodied agents, with a particular emphasis on the intricate task of selecting substitute objects. We introduce the CommonSense Object Affordance Task…

人工智能 · 计算机科学 2024-10-24 Ayush Agrawal , Raghav Prabhakar , Anirudh Goyal , Dianbo Liu

Controlling embodied agents with many actuated degrees of freedom is a challenging task. We propose a method that can discover and interpolate between context dependent high-level actions or body-affordances. These provide an abstract,…

人工智能 · 计算机科学 2017-08-16 Nicholas Guttenberg , Martin Biehl , Ryota Kanai

Full 3D estimation of human pose from a single image remains a challenging task despite many recent advances. In this paper, we explore the hypothesis that strong prior information about scene geometry can be used to improve pose estimation…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Zhe Wang , Liyan Chen , Shaurya Rathore , Daeyun Shin , Charless Fowlkes

Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either relies on robot data…

机器人学 · 计算机科学 2026-05-05 Yifan Han , Jianxiang Liu , Haoyu Zhang , Yuqi Gu , Yunhan Guo , Wenzhao Lian

We consider the task of learning to estimate human pose in still images. In order to avoid the high cost of full supervision, we propose to use a diverse data set, which consists of two types of annotations: (i) a small number of images are…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Aditya Arun , C. V. Jawahar , M. Pawan Kumar

The ability to autonomously explore and navigate a physical space is a fundamental requirement for virtually any mobile autonomous agent, from household robotic vacuums to autonomous vehicles. Traditional SLAM-based approaches for…

机器人学 · 计算机科学 2020-02-18 William Qi , Ravi Teja Mullapudi , Saurabh Gupta , Deva Ramanan

Does having visual priors (e.g. the ability to detect objects) facilitate learning to perform vision-based manipulation (e.g. picking up objects)? We study this problem under the framework of transfer learning, where the model is first…

机器人学 · 计算机科学 2021-07-02 Lin Yen-Chen , Andy Zeng , Shuran Song , Phillip Isola , Tsung-Yi Lin

We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object…

机器学习 · 计算机科学 2012-07-03 Yun Jiang , Marcus Lim , Ashutosh Saxena

How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambiguities by leveraging contextual knowledge -- such as affordances, where an object's shape and…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Naru Suzuki , Takehiko Ohkawa , Tatsuro Banno , Jihyun Lee , Ryosuke Furuta , Yoichi Sato

Grasping occluded objects in cluttered environments is an essential component in complex robotic manipulation tasks. In this paper, we introduce an AffordanCE-driven Next-Best-View planning policy (ACE-NBV) that tries to find a feasible…

In this work, we tackle one-shot visual search of object parts. Given a single reference image of an object with annotated affordance regions, we segment semantically corresponding parts within a target scene. We propose AffCorrs, an…

计算机视觉与模式识别 · 计算机科学 2022-09-19 Denis Hadjivelichkov , Sicelukwanda Zwane , Marc Peter Deisenroth , Lourdes Agapito , Dimitrios Kanoulas

Autonomous agents must often detect affordances: the set of behaviors enabled by a situation. Affordance detection is particularly helpful in domains with large action spaces, allowing the agent to prune its search space by avoiding futile…

人工智能 · 计算机科学 2018-09-03 Nancy Fulda , Daniel Ricks , Ben Murdoch , David Wingate

Lifelong user behavior sequences are crucial for capturing user interests and predicting user responses in modern recommendation systems. A two-stage paradigm is typically adopted to handle these long sequences: a subset of relevant…

信息检索 · 计算机科学 2025-03-27 Ningya Feng , Junwei Pan , Jialong Wu , Baixu Chen , Ximei Wang , Qian Li , Xian Hu , Jie Jiang , Mingsheng Long

This paper explores the grounding issue regarding multimodal semantic representation from a computational cognitive-linguistic view. We annotate images from the Flickr30k dataset with five perceptual properties: Affordance, Perceptual…

计算与语言 · 计算机科学 2023-10-25 Pin-Er Chen , Po-Ya Angela Wang , Hsin-Yu Chou , Yu-Hsiang Tseng , Shu-Kai Hsieh

The choice of a grasp plays a critical role in the success of downstream manipulation tasks. Consider a task of placing an object in a cluttered scene; the majority of possible grasps may not be suitable for the desired placement. In this…

机器人学 · 计算机科学 2023-04-11 Zhanpeng He , Nikhil Chavan-Dafle , Jinwook Huh , Shuran Song , Volkan Isler

Understanding and manipulating deformable objects (e.g., ropes and fabrics) is an essential yet challenging task with broad applications. Difficulties come from complex states and dynamics, diverse configurations and high-dimensional action…

计算机视觉与模式识别 · 计算机科学 2023-07-24 Ruihai Wu , Chuanruo Ning , Hao Dong

In order to *generalize* to various tasks in the wild, robotic agents will need a suitable representation (i.e., vision network) that enables the robot to predict optimal actions given high dimensional vision inputs. However, learning such…

机器人学 · 计算机科学 2024-07-29 Mohan Kumar Srirama , Sudeep Dasari , Shikhar Bahl , Abhinav Gupta

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Tushar Nagarajan , Yanghao Li , Christoph Feichtenhofer , Kristen Grauman