English
Related papers

Related papers: Intention Action Anticipation Model with Guide-Fee…

200 papers

Anticipating human intention by observing one's actions has many applications. For instance, picking up a cellphone, then a charger (actions) implies that one wants to charge the cellphone (intention). By anticipating the intention, an…

Computer Vision and Pattern Recognition · Computer Science 2017-10-23 Tz-Ying Wu , Ting-An Chien , Cheng-Sheng Chan , Chan-Wei Hu , Min Sun

Multi-person motion prediction is an emerging and intricate task with broad real-world applications. Unlike single person motion prediction, it considers not just the skeleton structures or human trajectories but also the interactions…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Kehua Qu , Rui Ding , Jin Tang

Interactive assistance systems typically provide feedback after an action has been completed, supporting error recovery but not preventing the error itself. We present TRAFA, a real-time predictive feedback system for procedural tasks that…

Human-Computer Interaction · Computer Science 2026-05-26 Sassan Mokhtar , Lars Doorenbos , Fatemeh Jabbari , Marius Bock , Dominik Bach , Juergen Gall

Visual Commonsense Reasoning (VCR) remains a significant yet challenging research problem in the realm of visual reasoning. A VCR model generally aims at answering a textual question regarding an image, followed by the rationale prediction…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Zhenyang Li , Yangyang Guo , Kejie Wang , Fan Liu , Liqiang Nie , Mohan Kankanhalli

Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recovering either hand or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Yufei Ye , Jiaman Li , Ryan Rong , C. Karen Liu

A promising effective human-robot interaction in assistive robotic systems is gaze-based control. However, current gaze-based assistive systems mainly help users with basic grasping actions, offering limited support. Moreover, the…

Robotics · Computer Science 2025-08-20 Zejia Zhang , Bo Yang , Xinxing Chen , Weizhuang Shi , Haoyuan Wang , Wei Luo , Jian Huang

Within this work, we explore intention inference for user actions in the context of a handheld robot setup. Handheld robots share the shape and properties of handheld tools while being able to process task information and aid manipulation.…

Robotics · Computer Science 2019-03-21 Janis Stolzenwald , Walterio W. Mayol-Cuevas

Recently, deception detection on human videos is an eye-catching techniques and can serve lots applications. AI model in this domain demonstrates the high accuracy, but AI tends to be a non-interpretable black box. We introduce an…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Shun-Wen Hsiao , Cheng-Yuan Sun

Human trajectory forecasting is important for intelligent multimedia systems operating in visually complex environments, such as autonomous driving and crowd surveillance. Although Conditional Flow Matching (CFM) has shown strong ability in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Xuepeng Jing , Wenhuan Lu , Hao Meng , Zhizhi Yu , Jianguo Wei

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

Multimedia · Computer Science 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

Ego-to-exo video generation refers to generating the corresponding exocentric video according to the egocentric video, providing valuable applications in AR/VR and embodied AI. Benefiting from advancements in diffusion model techniques,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Hongchen Luo , Kai Zhu , Wei Zhai , Yang Cao

Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. We study whether short-horizon, task-consistent future videos can provide useful…

Robotics · Computer Science 2026-05-29 Mohammad Khoshnazar , Andrew Melnik , Michael Beetz

While deep reinforcement learning (RL) agents outperform humans on an increasing number of tasks, training them requires data equivalent to decades of human gameplay. Recent hierarchical RL methods have increased sample efficiency by…

Machine Learning · Computer Science 2023-06-21 Anna Penzkofer , Simon Schaefer , Florian Strohm , Mihai Bâce , Stefan Leutenegger , Andreas Bulling

Action anticipation involves predicting future actions having observed the initial portion of a video. Typically, the observed video is processed as a whole to obtain a video-level representation of the ongoing activity in the video, which…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Megha Nawhal , Akash Abdu Jyothi , Greg Mori

Egocentric action anticipation is the task of predicting the future actions a camera wearer will likely perform based on past video observations. While in a real-world system it is fundamental to output such predictions before the action…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Antonino Furnari , Giovanni Maria Farinella

To engender safe and efficient human-robot collaboration, it is critical to generate high-fidelity predictions of human behavior. The challenges in making accurate predictions lie in the stochasticity and heterogeneity in human behaviors.…

Robotics · Computer Science 2019-09-12 Abulikemu Abuduweili , Siyan Li , Changliu Liu

Trajectory prediction is of significant importance in computer vision. Accurate pedestrian trajectory prediction benefits autonomous vehicles and robots in planning their motion. Pedestrians' trajectories are greatly influenced by their…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Pengqian Han , Jiamou Liu , Jialing He , Zeyu Zhang , Song Yang , Yanni Tang

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three…

Computer Vision and Pattern Recognition · Computer Science 2019-08-23 Evangelos Kazakos , Arsha Nagrani , Andrew Zisserman , Dima Damen

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

Derived from rapid advances in computer vision and machine learning, video analysis tasks have been moving from inferring the present state to predicting the future state. Vision-based action recognition and prediction from videos are such…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Yu Kong , Yun Fu