中文
相关论文

相关论文: Visual-Geometric Collaborative Guidance for Afford…

200 篇论文

We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and…

计算机视觉与模式识别 · 计算机科学 2018-03-06 Thanh-Toan Do , Anh Nguyen , Ian Reid

Interactive object selection is a very important research problem and has many applications. Previous algorithms require substantial user interactions to estimate the foreground and background distributions. In this paper, we present a…

计算机视觉与模式识别 · 计算机科学 2016-03-15 Ning Xu , Brian Price , Scott Cohen , Jimei Yang , Thomas Huang

Affordance grounding aims to localize the interaction regions for the manipulated objects in the scene image according to given instructions. A critical challenge in affordance grounding is that the embodied agent should understand human…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Changmao Chen , Yuren Cong , Zhen Kan

Interactive reinforcement learning (IRL) extends traditional reinforcement learning (RL) by allowing an agent to interact with parent-like trainers during a task. In this paper, we present an IRL approach using dynamic audio-visual input in…

人工智能 · 计算机科学 2018-07-27 Francisco Cruz , German I. Parisi , Stefan Wermter

To enable remote Virtual Reality (VR) and Augmented Reality (AR) clients to collaborate as if they were in the same space during Mixed Reality (MR) telepresence, it is essential to overcome spatial heterogeneity and generate a unified…

新兴技术 · 计算机科学 2024-08-09 Dooyoung Kim , Seonji Kim , Selin Choi , Woontack Woo

Motivated by the intuitive understanding humans have about the space of possible interactions, and the ease with which they can generalize this understanding to previously unseen scenes, we develop an approach for learning visual…

机器人学 · 计算机科学 2023-05-30 Homanga Bharadhwaj , Abhinav Gupta , Shubham Tulsiani

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Yan Zhao , Ruihai Wu , Zhehuan Chen , Yourong Zhang , Qingnan Fan , Kaichun Mo , Hao Dong

Can a video generation model be repurposed as an interactive world simulator? We explore the affordance perception potential of text-to-video models by teaching them to predict human-environment interaction. Given a scene image and a prompt…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Mengyi Shan , Zecheng He , Haoyu Ma , Felix Juefei-Xu , Peizhao Zhang , Tingbo Hou , Ching-Yao Chuang

Interaction in virtual reality (VR) environments is essential to achieve a pleasant and immersive experience. Most of the currently existing VR applications, lack of robust object grasping and manipulation, which are the cornerstone of…

Service robots are expected to autonomously and efficiently work in human-centric environments. For this type of robots, object perception and manipulation are challenging tasks due to need for accurate and real-time response. This paper…

机器人学 · 计算机科学 2019-04-05 S. Hamidreza Kasaei , Nima Shafii , Luis Seabra Lopes , Ana Maria Tome

Currently, task-oriented grasp detection approaches are mostly based on pixel-level affordance detection and semantic segmentation. These pixel-level approaches heavily rely on the accuracy of a 2D affordance mask, and the generated grasp…

机器人学 · 计算机科学 2022-10-18 Wenkai Chen , Hongzhuo Liang , Zhaopeng Chen , Fuchun Sun , Jianwei Zhang

In this work, we focus on the task of weakly supervised affordance grounding, where a model is trained to identify affordance regions on objects using human-object interaction images and egocentric object images without dense labels.…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Peiran Xu , Yadong Mu

How do we know that a kitchen is a kitchen by looking? Relatively little is known about how we conceptualize and categorize different visual environments. Traditional models of visual perception posit that scene categorization is achieved…

神经元与认知 · 定量生物学 2014-11-20 Michelle R. Greene , Christopher Baldassano , Andre Esteva , Diane M. Beck , Li Fei-Fei

For robots to exhibit a high level of intelligence in the real world, they must be able to assess objects for which they have no prior knowledge. Therefore, it is crucial for robots to perceive object affordances by reasoning about physical…

机器人学 · 计算机科学 2020-04-09 Hongtao Wu , Deven Misra , Gregory S. Chirikjian

This paper proposes an attributable visual similarity learning (AVSL) framework for a more accurate and explainable similarity measure between images. Most existing similarity learning methods exacerbate the unexplainability by mapping each…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Borui Zhang , Wenzhao Zheng , Jie Zhou , Jiwen Lu

We introduce One-shot Open Affordance Learning (OOAL), where a model is trained with just one example per base object category, but is expected to identify novel objects and affordances. While vision-language models excel at recognizing…

计算机视觉与模式识别 · 计算机科学 2023-11-30 Gen Li , Deqing Sun , Laura Sevilla-Lara , Varun Jampani

Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and recognize their interactions. While recent two-stage approaches have made significant progress, they still face challenges due to incomplete…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Zhehao Li , Yucheng Qian , Chong Wang , Yinghao Lu , Zhihao Yang , Jiafei Wu

Enabling humans and robots to collaborate effectively requires purposeful communication and an understanding of each other's affordances. Prior work in human-robot collaboration has incorporated knowledge of human affordances, i.e., their…

机器人学 · 计算机科学 2023-12-22 Drake Moore , Mark Zolotas , Taskin Padir

Many active learning and search approaches are intractable for large-scale industrial settings with billions of unlabeled examples. Existing approaches search globally for the optimal examples to label, scaling linearly or even…

Enabling robotic manipulation that generalizes to out-of-distribution scenes is a crucial step toward open-world embodied intelligence. For human beings, this ability is rooted in the understanding of semantic correspondence among objects,…

机器人学 · 计算机科学 2024-01-17 Yuanchen Ju , Kaizhe Hu , Guowei Zhang , Gu Zhang , Mingrun Jiang , Huazhe Xu