English
Related papers

Related papers: Anticipating Next Active Objects for Egocentric Vi…

200 papers

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person…

Action anticipation involves predicting future actions having observed the initial portion of a video. Typically, the observed video is processed as a whole to obtain a video-level representation of the ongoing activity in the video, which…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Megha Nawhal , Akash Abdu Jyothi , Greg Mori

This paper deals with the problem of localizing objects in image and video datasets from visual exemplars. In particular, we focus on the challenging problem of egocentric visual query localization. We first identify grave implicit biases…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Mengmeng Xu , Yanghao Li , Cheng-Yang Fu , Bernard Ghanem , Tao Xiang , Juan-Manuel Perez-Rua

Predicting future human behavior from egocentric videos is a challenging but critical task for human intention understanding. Existing methods for forecasting 2D hand positions rely on visual representations and mainly focus on hand-object…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Masashi Hatano , Ryo Hachiuma , Hideo Saito

The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In this work, we introduce the challenging problem of predicting…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Boxiao Pan , Adam W. Harley , C. Karen Liu , Leonidas J. Guibas

In this work we employ multitask learning to capitalize on the structure that exists in related supervised tasks to train complex neural networks. It allows training a network for multiple objectives in parallel, in order to improve…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas Noldus , Remco Veltkamp

Speed-control forecasting, a challenging problem in driver behavior analysis, aims to predict the future actions of a driver in controlling vehicle speed such as braking or acceleration. In this paper, we try to address this challenge…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Yichen Ding , Ziming Zhang , Yanhua Li , Xun Zhou

In this paper, we address the problem of forecasting the trajectory of an egocentric camera wearer (ego-person) in crowded spaces. The trajectory forecasting ability learned from the data of different camera wearers walking around in the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-08 Jianing Qiu , Lipeng Chen , Xiao Gu , Frank P. -W. Lo , Ya-Yen Tsai , Jiankai Sun , Jiaqi Liu , Benny Lo

Learning to infer labels in an open world, i.e., in an environment where the target ``labels'' are unknown, is an important characteristic for achieving autonomy. Foundation models, pre-trained on enormous amounts of data, have shown…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Sanjoy Kundu , Shubham Trehan , Sathyanarayanan N. Aakur

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

In this paper, we address the problem of short-term action anticipation, i.e., we want to predict an upcoming action one second before it happens. We propose to harness high-level intent information to anticipate actions that will take…

Computer Vision and Pattern Recognition · Computer Science 2023-06-28 Olga Zatsarynna , Juergen Gall

Understanding the camera wearer's activity is central to egocentric vision, yet one key facet of that activity is inherently invisible to the camera--the wearer's body pose. Prior work focuses on estimating the pose of hands and arms when…

Computer Vision and Pattern Recognition · Computer Science 2016-03-28 Hao Jiang , Kristen Grauman

Temporal action localization aims to predict the boundary and category of each action instance in untrimmed long videos. Most of previous methods based on anchors or proposals neglect the global-local context interaction in entire video…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Yizheng Ouyang , Tianjin Zhang , Weibo Gu , Hongfa Wang

Locating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer. Current temporal action localization methods typically rely…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Erhang Zhang , Junyi Ma , Yin-Dong Zheng , Yixuan Zhou , Hesheng Wang

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task remains a highly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jaeyoung Choi , Hyeondong Kim , Yujin Kim , Daehee Park

We present a unified framework for understanding 3D hand and object interactions in raw image sequences from egocentric RGB cameras. Given a single RGB image, our model jointly estimates the 3D hand and object poses, models their…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Bugra Tekin , Federica Bogo , Marc Pollefeys

The task of predicting future actions from a video is crucial for a real-world agent interacting with others. When anticipating actions in the distant future, we humans typically consider long-term relations over the whole sequence of…

Computer Vision and Pattern Recognition · Computer Science 2022-05-30 Dayoung Gong , Joonseok Lee , Manjin Kim , Seong Jong Ha , Minsu Cho

In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoint changes in egocentric videos cause significant object…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Joohyun Chang , Soyeon Hong , Hyogun Lee , Seong Jong Ha , Dongho Lee , Seong Tae Kim , Jinwoo Choi

To have a robot actively supporting a human during a collaborative task, it is crucial that robots are able to identify the current action in order to predict the next one. Common approaches make use of high-level knowledge, such as object…

Robotics · Computer Science 2017-03-08 Markus Eich , Sareh Shirazi , Gordon Wyeth

Human intention detection with hand motion prediction is critical to drive the upper-extremity assistive robots in neurorehabilitation applications. However, the traditional methods relying on physiological signal measurement are…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Yufei He , Xucong Zhang , Arno H. A. Stienen