English
Related papers

Related papers: Gaze-Guided Graph Neural Network for Action Antici…

200 papers

Modeling and automatically recognizing surgical activities are fundamental steps toward automation in surgery and play important roles in providing timely feedback to surgeons. Accurately recognizing surgical activities in video poses a…

Image and Video Processing · Electrical Eng. & Systems 2022-11-15 Abdishakour Awale , Duygu Sarikaya

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Taiying Peng , Jiacheng Hua , Miao Liu , Feng Lu

Gaze is a powerful form of non-verbal communication and social interaction that humans develop from an early age. As such, modeling this behavior is an important task that can benefit a broad set of application domains ranging from robotics…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Samy Tafasca , Anshul Gupta , Jean-Marc Odobez

This paper addresses a new problem of understanding human gaze communication in social videos from both atomic-level and event-level, which is significant for studying human social interactions. To tackle this novel and challenging problem,…

Computer Vision and Pattern Recognition · Computer Science 2019-09-06 Lifeng Fan , Wenguan Wang , Siyuan Huang , Xinyu Tang , Song-Chun Zhu

In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context. Our assumption is that in the procedure of performing a manipulation task, what a person is doing…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Yifei Huang , Zhenqiang Li , Minjie Cai , Yoichi Sato

Humans directly completing tasks in dangerous or hazardous conditions is not always possible where these tasks are increasingly be performed remotely by teleoperated robots. However, teleoperation is difficult since the operator feels a…

Robotics · Computer Science 2025-04-07 Jeremy D. Webb , Michael Bowman , Songpo Li , Xiaoli Zhang

Learning new skills by observing humans' behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans' decision-making processes, focusing on learning a model to plan goal-directed actions in…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Jing Bi , Jiebo Luo , Chenliang Xu

Predicting the next action that a human is most likely to perform is key to human-AI collaboration and has consequently attracted increasing research interests in recent years. An important factor for next action prediction are human…

Human-Computer Interaction · Computer Science 2024-03-26 Lei Shi , Paul-Christian Bürkner , Andreas Bulling

For machines to effectively assist humans in challenging visual search tasks, they must differentiate whether a human is simply glancing into a scene (navigational intent) or searching for a target object (informational intent). Previous…

Human-Computer Interaction · Computer Science 2025-08-05 Mansi Sharma , Shuang Chen , Philipp Müller , Maurice Rekrut , Antonio Krüger

We present GazeMotion, a novel method for human motion forecasting that combines information on past human poses with human eye gaze. Inspired by evidence from behavioural sciences showing that human eye and body movements are closely…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Zhiming Hu , Syn Schmitt , Daniel Haeufle , Andreas Bulling

We address goal-based imitation learning, where the aim is to output the symbolic goal from a third-person video demonstration. This enables the robot to plan for execution and reproduce the same goal in a completely different environment.…

The action anticipation task refers to predicting what action will happen based on observed videos, which requires the model to have a strong ability to summarize the present and then reason about the future. Experience and common sense…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Xin Liu , Chao Hao , Zitong Yu , Huanjing Yue , Jingyu Yang

Most human behaviors consist of multiple parts, steps, or subtasks. These structures guide our action planning and execution, but when we observe others, the latent structure of their actions is typically unobservable, and must be inferred…

Artificial Intelligence · Computer Science 2018-09-28 Ryo Nakahashi , Chris L. Baker , Joshua B. Tenenbaum

Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates on fully observed videos, action anticipation must handle…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Seulgi Kim , Ghazal Kaviani , Mohit Prabhushankar , Ghassan AlRegib

We tackle the problem of predicting saliency maps for videos of dynamic scenes. We note that the accuracy of the maps reconstructed from the gaze data of a fixed number of observers varies with the frame, as it depends on the content of the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Ekta Prashnani , Orazio Gallo , Joohwan Kim , Josef Spjut , Pradeep Sen , Iuri Frosio

The primary objective of the dataset is to provide a better understanding of the coupling between human actions and gaze in a shared working environment with a cobot, with the aim of signifcantly enhancing the effciency and safety of…

Robotics · Computer Science 2025-03-17 Maxence Grand , Damien Pellier , Francis Jambon

Emotion recognition,as a step toward mind reading,seeks to infer internal states from external cues.Most existing methods rely on explicit signals-such as facial expressions,speech,or gestures-that reflect only bodily responses and overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Mengke Song , Yuge Xie , Qi Cui , Luming Li , Xinyu Liu , Guotao Wang , Chenglizhao Chen , Shanchen Pang

We present an attention based visual analysis framework to compute grasp-relevant information in order to guide grasp planning using a multi-fingered robotic hand. Our approach uses a computational visual attention model to locate regions…

Robotics · Computer Science 2018-09-13 Zhen Deng , Ge Gao , Simone Frintrop , Jianwei Zhang

Due to the recent outbreak of COVID-19, many classes, exams, and meetings have been conducted non-face-to-face. However, the foundation for video conferencing solutions is still insufficient. So this technology has become an important…

Computer Vision and Pattern Recognition · Computer Science 2021-07-15 Suneung-Kim , Seong-Whan Lee

Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realistic applications such as Augmented Reality (AR) glasses. While prior streaming benchmarks evaluate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Daeun Lee , Subhojyoti Mukherjee , Branislav Kveton , Ryan A. Rossi , Viet Dac Lai , Seunghyun Yoon , Trung Bui , Franck Dernoncourt , Mohit Bansal