中文
相关论文

相关论文: PlayerOne: Egocentric World Simulator

200 篇论文

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Weikun Peng , Denys Iliash , Manolis Savva

We argue that 3-D first-person video games are a challenging environment for real-time multi-modal reasoning. We first describe our dataset of human game-play, collected across a large variety of 3-D first-person games, which is both…

机器学习 · 计算机科学 2025-10-21 Yuguang Yue , Irakli Salia , Samuel Hunt , Christopher Green , Wenzhe Shi , Jonathan J Hunt

World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Xiaofeng Wang , Zheng Zhu , Guan Huang , Xinze Chen , Jiagang Zhu , Jiwen Lu

Egocentric videos provide valuable insights into human interactions with the physical world, which has sparked growing interest in the computer vision and robotics communities. A critical challenge in fully understanding the geometry and…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Chengbo Yuan , Geng Chen , Li Yi , Yang Gao

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video…

人工智能 · 计算机科学 2026-05-14 Qinchuan Cheng , Zhantao Gong , Pengzhan Sun , Angela Yao , Xulei Yang , Shijie Li

Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requires extensive modeling and manual setup, and more…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Quankai Gao , Jiawei Yang , Qiangeng Xu , Le Chen , Yue Wang

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Yutong Bai , Danny Tran , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik

Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans,…

Egocentric videos can bring a lot of information about how humans perceive the world and interact with the environment, which can be beneficial for the analysis of human behaviour. The research in egocentric video analysis is developing…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Ivan Rodin , Antonino Furnari , Dimitrios Mavroedis , Giovanni Maria Farinella

We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Mi Luo , Zihui Xue , Alex Dimakis , Kristen Grauman

Motion-controllable video generation is crucial for egocentric applications in virtual reality and embodied AI. However, existing methods often struggle to achieve 3D-consistent fine-grained hand articulation. By adopting on 2D trajectories…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Chenyangguang Zhang , Botao Ye , Boqi Chen , Alexandros Delitzas , Fangjinhua Wang , Marc Pollefeys , Xi Wang

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Gen Li , Kaifeng Zhao , Siwei Zhang , Xiaozhong Lyu , Mihai Dusmanu , Yan Zhang , Marc Pollefeys , Siyu Tang

Recently, there has been a growing interest in wearable sensors which provides new research perspectives for 360 {\deg} video analysis. However, the lack of 360 {\deg} datasets in literature hinders the research in this field. To bridge…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Keshav Bhandari , Mario A. DeLaGarza , Ziliang Zong , Hugo Latapie , Yan Yan

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid egomotion and…

The advancement of robot learning is currently hindered by the scarcity of large-scale, high-quality datasets. While established data collection methods such as teleoperation and universal manipulation interfaces dominate current datasets,…

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Junho Park , Andrew Sangwoo Ye , Taein Kwon

Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recovering either hand or…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Yufei Ye , Jiaman Li , Ryan Rong , C. Karen Liu