中文
相关论文

相关论文: PlayerOne: Egocentric World Simulator

200 篇论文

We present GameNGen, the first game engine powered entirely by a neural model that also enables real-time interaction with a complex environment over long trajectories at high quality. When trained on the classic game DOOM, GameNGen…

机器学习 · 计算机科学 2025-04-25 Dani Valevski , Yaniv Leviathan , Moab Arar , Shlomi Fruchter

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video generation with evolving…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Jiahao Wang , Luoxin Ye , TaiMing Lu , Junfei Xiao , Jiahan Zhang , Yuxiang Guo , Xijun Liu , Rama Chellappa , Cheng Peng , Alan Yuille , Jieneng Chen

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in egocentric settings and…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Francesco Ragusa , Antonino Furnari , Salvatore Livatino , Giovanni Maria Farinella

Egocentric, or first-person vision which became popular in recent years with an emerge in wearable technology, is different than exocentric (third-person) vision in some distinguishable ways, one of which being that the camera wearer is…

计算机视觉与模式识别 · 计算机科学 2016-10-11 Jessica Finocchiaro , Aisha Urooj Khan , Ali Borji

Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a real-time framework for infinite, 3D-consistent video…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Ke Zhang , Yiqun Mei , Jiacong Xu , Vishal M. Patel

Video generation models are increasingly used as world simulators for storytelling, simulation, and embodied AI. As these models advance, a key question arises: do generated videos obey the physical laws of the real world? Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Qin Zhang , Peiyu Jing , Hong-Xing Yu , Fangqiang Ding , Fan Nie , Weimin Wang , Yilun Du , James Zou , Jiajun Wu , Bing Shuai

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Autonomous driving systems struggle with complex scenarios due to limited access to diverse, extensive, and out-of-distribution driving data which are critical for safe navigation. World models offer a promising solution to this challenge;…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Xi Guo , Chenjing Ding , Haoxuan Dou , Xin Zhang , Weixuan Tang , Wei Wu

Recently, video-based world models that learn to simulate the dynamics have gained increasing attention in robot learning. However, current approaches primarily emphasize visual generative quality while overlooking physical fidelity,…

机器人学 · 计算机科学 2026-01-21 Baorui Peng , Wenyao Zhang , Liang Xu , Zekun Qi , Jiazhao Zhang , Hongsi Liu , Wenjun Zeng , Xin Jin

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Siyuan Zhou , Yilun Du , Yuncong Yang , Lei Han , Peihao Chen , Dit-Yan Yeung , Chuang Gan

Video game playing is an extremely structured domain where algorithmic decision-making can be tested without adverse real-world consequences. While prevailing methods rely on image inputs to avoid the problem of hand-crafting state space…

机器学习 · 计算机科学 2024-09-24 Abhishek Jaiswal , Nisheeth Srivastava

We introduce Matrix-Game, an interactive world foundation model for controllable game world generation. Matrix-Game is trained using a two-stage pipeline that first performs large-scale unlabeled pretraining for environment understanding,…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Yifan Zhang , Chunli Peng , Boyang Wang , Puyi Wang , Qingcheng Zhu , Fei Kang , Biao Jiang , Zedong Gao , Eric Li , Yang Liu , Yahui Zhou

It has been challenging to find ways to educate people to have better environmental consciousness. In some cases, people do not know what the right behaviors are to protect the environment. Game engine has been used in the AEC industry for…

计算机与社会 · 计算机科学 2023-04-24 Jing Du

In video games, players' perception of the game world and related information depends on their or the game designer's choice of a virtual camera model. In this paper, we attempt to answer the research question of whether it is possible to…

多媒体 · 计算机科学 2021-09-09 Markos Naftis , George Tsatiris , Kostas Karpouzis

Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle with a lack of spatial persistence and insufficient visual…

Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner…

计算机视觉与模式识别 · 计算机科学 2022-08-17 Siwei Zhang , Qianli Ma , Yan Zhang , Zhiyin Qian , Taein Kwon , Marc Pollefeys , Federica Bogo , Siyu Tang

We propose a method for object-aware 3D egocentric pose estimation that tightly integrates kinematics modeling, dynamics modeling, and scene object information. Unlike prior kinematics or dynamics-based approaches where the two components…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Zhengyi Luo , Ryo Hachiuma , Ye Yuan , Kris Kitani

Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective distortion, and noisy ego-motion. Existing methods rely on…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Bohan Zhou , Yi Zhan , Zhongbin Zhang , Zongqing Lu

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

计算机视觉与模式识别 · 计算机科学 2016-08-16 Shervin Ardeshir , Ali Borji