中文
相关论文

相关论文: Gaze-Guided Graph Neural Network for Action Antici…

200 篇论文

In this work we employ multitask learning to capitalize on the structure that exists in related supervised tasks to train complex neural networks. It allows training a network for multiple objectives in parallel, in order to improve…

计算机视觉与模式识别 · 计算机科学 2019-09-17 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas Noldus , Remco Veltkamp

Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentric action anticipation task, which predicts future action…

计算机视觉与模式识别 · 计算机科学 2021-01-20 Yu Wu , Linchao Zhu , Xiaohan Wang , Yi Yang , Fei Wu

Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current…

计算机视觉与模式识别 · 计算机科学 2019-12-17 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works have explored leveraging human data to alleviate this…

机器人学 · 计算机科学 2026-05-01 Chengyang Li , Kaiyi Xiong , Yuan Xu , Lei Qian , Yizhou Wang , Wentao Zhu

We present a novel approach for the visual prediction of human-object interactions in videos. Rather than forecasting the human and object motion or the future hand-object contact points, we aim at predicting (a)the class of the on-going…

计算机视觉与模式识别 · 计算机科学 2022-09-13 Victoria Manousaki , Konstantinos Papoutsakis , Antonis Argyros

Mixed Reality (MR) interfaces increasingly rely on gaze for interaction , yet distinguishing visual attention from intentional action remains difficult, leading to the Midas Touch problem. Existing solutions require explicit confirmations,…

人机交互 · 计算机科学 2026-01-28 Francesco Chiossi , Elnur Imamaliyev , Martin Bleichner , Sven Mayer

We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. To facilitate our research, we first introduce the EGTEA Gaze+ dataset. Our dataset…

计算机视觉与模式识别 · 计算机科学 2020-11-03 Yin Li , Miao Liu , James M. Rehg

Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to…

机器人学 · 计算机科学 2026-05-29 Kuangji Zuo , Gen Li , Bofan Lyu , Yanshuo Lu , Boyu Ma , Shijia Han , Xinyu Zhou , Xichen Yuan , Chuhao Zhou , Jiaqi Bai , Geng Li , Jianfei Yang

This paper explores the estimation of user attention in the setting of a cooperative handheld robot: a robot designed to behave as a handheld tool but that has levels of task knowledge. We use a tool-mounted gaze tracking system, which,…

机器人学 · 计算机科学 2018-10-16 Janis Stolzenwald , Walterio W. Mayol-Cuevas

This paper introduces a new and challenging Hidden Intention Discovery (HID) task. Unlike existing intention recognition tasks, which are based on obvious visual representations to identify common intentions for normal behavior, HID focuses…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Zhuo Zhou , Wenxuan Liu , Danni Xu , Zheng Wang , Jian Zhao

This paper presents the selective use of eye-gaze information in learning human actions in Atari games. Vast evidence suggests that our eye movement convey a wealth of information about the direction of our attention and mental states and…

机器学习 · 计算机科学 2020-12-08 Chaitanya Thammineni , Hemanth Manjunatha , Ehsan T. Esfahani

The goal of visual analytics is to create a symbiosis between human and computer by leveraging their unique strengths. While this model has demonstrated immense success, we are yet to realize the full potential of such a human-computer…

人机交互 · 计算机科学 2018-09-27 Ran Wan , Roman Garnett , Alvitta Ottley

Video understanding is to recognize and classify different actions or activities appearing in the video. A lot of previous work, such as video captioning, has shown promising performance in producing general video understanding. However, it…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Zijian Kuang , Xinran Tie

Anticipating future actions based on spatiotemporal observations is essential in video understanding and predictive computer vision. Moreover, a model capable of anticipating the future has important applications, it can benefit…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Tsung-Ming Tai , Giuseppe Fiameni , Cheng-Kuang Lee , Simon See , Oswald Lanz

Egocentric perception has grown rapidly with the advent of immersive computing devices. Human gaze prediction is an important problem in analyzing egocentric videos and has primarily been tackled through either saliency-based modeling or…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Sathyanarayanan N. Aakur , Arunkumar Bagavathi

Visual attention plays a critical role when our visual system executes active visual tasks by interacting with the physical scene. However, how to encode the visual object relationship in the psychological world of our brain deserves to be…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Kai-Fu Yang , Yong-Jie Li

Though action recognition in videos has achieved great success recently, it remains a challenging task due to the massive computational cost. Designing lightweight networks is a possible solution, but it may degrade the recognition…

计算机视觉与模式识别 · 计算机科学 2020-02-11 Wenhao Wu , Dongliang He , Xiao Tan , Shifeng Chen , Yi Yang , Shilei Wen

People are proficient at communicating their intentions in order to avoid conflicts when navigating in narrow, crowded environments. In many situations mobile robots lack both the ability to interpret human intentions and the ability to…

机器人学 · 计算机科学 2019-11-07 Justin Hart , Reuth Mirsky , Stone Tejeda , Bonny Mahajan , Jamin Goo , Kathryn Baldauf , Sydney Owen , Peter Stone

Action prediction aims to infer the forthcoming human action with partially-observed videos, which is a challenging task due to the limited information underlying early observations. Existing methods mainly adopt a reconstruction strategy…

计算机视觉与模式识别 · 计算机科学 2021-12-21 Zhiqiang Tao , Yue Bai , Handong Zhao , Sheng Li , Yu Kong , Yun Fu

This paper presents a teleoperation system that includes robot perception and intent prediction from hand gestures. The perception module identifies the objects present in the robot workspace and the intent prediction module which object…

机器人学 · 计算机科学 2021-07-06 Yoojin Oh , Marc Toussaint , Jim Mainprice