English
Related papers

Related papers: EgoReAct: Egocentric Video-Driven 3D Human Reactio…

200 papers

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

Robotics · Computer Science 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

Predicting future human behavior from egocentric videos is a challenging but critical task for human intention understanding. Existing methods for forecasting 2D hand positions rely on visual representations and mainly focus on hand-object…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Masashi Hatano , Ryo Hachiuma , Hideo Saito

Multi-view egocentric dynamic scene reconstruction holds significant research value for applications in holographic documentation of social interactions. However, existing reconstruction datasets focus on static multi-view or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Bate Li , Houqiang Zhong , Zhengxue Cheng , Qiang Hu , Qiang Wang , Li Song , Wenjun Zhang

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Yifei Huang , Guo Chen , Jilan Xu , Mingfang Zhang , Lijin Yang , Baoqi Pei , Hongjie Zhang , Lu Dong , Yali Wang , Limin Wang , Yu Qiao

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Jilan Xu , Yifei Huang , Baoqi Pei , Junlin Hou , Qingqiu Li , Guo Chen , Yuejie Zhang , Rui Feng , Weidi Xie

The advancement of robot learning is currently hindered by the scarcity of large-scale, high-quality datasets. While established data collection methods such as teleoperation and universal manipulation interfaces dominate current datasets,…

Intelligent assistance involves not only understanding but also action. Existing ego-centric video datasets contain rich annotations of the videos, but not of actions that an intelligent assistant could perform in the moment. To address…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Steven Abreu , Tiffany D. Do , Karan Ahuja , Eric J. Gonzalez , Lee Payne , Daniel McDuff , Mar Gonzalez-Franco

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Hui Li , Mingwang Xu , Yun Zhan , Shan Mu , Jiaye Li , Kaihui Cheng , Yuxuan Chen , Tan Chen , Mao Ye , Jingdong Wang , Siyu Zhu

Natural interaction with virtual objects in AR/VR environments makes for a smooth user experience. Gestures are a natural extension from real world to augmented space to achieve these interactions. Finding discriminating spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Tejo Chalasani , Jan Ondrej , Aljosa Smolic

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans'…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Danny Tran , Roberto Martín-Martín , Kristen Grauman

In recent years, we have seen the performance of video-based person Re-Identification (ReID) methods have improved considerably. However, most of the work in this area has dealt with videos acquired by fixed cameras with wider field of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Emrah Basaran , Yonatan Tariku Tesfaye , Mubarak Shah

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person…

Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Linxi Xie , Lisong C. Sun , Ashley Neall , Tong Wu , Shengqu Cai , Gordon Wetzstein

Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Junho Park , Andrew Sangwoo Ye , Taein Kwon

Learning an agent model that behaves like humans-capable of jointly perceiving the environment, predicting the future, and taking actions from a first-person perspective-is a fundamental challenge in computer vision. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Lu Chen , Yizhou Wang , Shixiang Tang , Qianhong Ma , Tong He , Wanli Ouyang , Xiaowei Zhou , Hujun Bao , Sida Peng

The proliferation of commercial egocentric devices offers a unique lens into human behavior, yet reconstructing full-body 3D motion remains difficult due to frequent self-occlusion and the 'out-of-sight' nature of the wearer's limbs. While…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Kyungwon Cho , Hanbyul Joo

"Looking for things" is a mundane but critical task we repeatedly carry on in our daily life. We introduce a method to develop a human character capable of searching for a randomly located target object in a detailed 3D scene using its…

Robotics · Computer Science 2021-09-16 Maks Sorokin , Wenhao Yu , Sehoon Ha , C. Karen Liu

Egocentric gestures are the most natural form of communication for humans to interact with wearable devices such as VR/AR helmets and glasses. A major issue in such scenarios for real-world applications is that may easily become necessary…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Zhengwei Wang , Qi She , Tejo Chalasani , Aljosa Smolic

Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Tatiana Tommasi , Giuseppe Averta

Generating realistic human-human interactions is a challenging task that requires not only high-quality individual body and hand motions, but also coherent coordination among all interactants. Due to limitations in available data and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Pablo Ruiz-Ponce , Sergio Escalera , José García-Rodríguez , Jiankang Deng , Rolandos Alexandros Potamias