English
Related papers

Related papers: MagicWorld: Towards Long-Horizon Stability for Int…

200 papers

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Ying Yang , Zhengyao Lv , Tianlin Pan , Haofan Wang , Binxin Yang , Hubery Yin , Chen Li , Ziwei Liu , Chenyang Si

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

Autoregressive video world models predict future visual observations conditioned on actions. While effective over short horizons, these models often struggle with long-horizon generation, as small prediction errors accumulate over time.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Junchao Huang , Ziyang Ye , Xinting Hu , Tianyu He , Guiyu Zhang , Shaoshuai Shi , Jiang Bian , Li Jiang

Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Zhiheng Liu , Xueqing Deng , Shoufa Chen , Angtian Wang , Qiushan Guo , Mingfei Han , Zeyue Xue , Mengzhao Chen , Ping Luo , Linjie Yang

Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implicitly assume that the world only evolves within the observer's…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zicheng Duan , Jiatong Xia , Zeyu Zhang , Wenbo Zhang , Gengze Zhou , Chenhui Gou , Yefei He , Feng Chen , Xinyu Zhang , Lingqiao Liu

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal context window sizes, these models often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Tong Wu , Shuai Yang , Ryan Po , Yinghao Xu , Ziwei Liu , Dahua Lin , Gordon Wetzstein

We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Hyeongju Mun , In-Hwan Jin , Sohyeong Kim , Kyeongbo Kong

This paper studies the human image animation task, which aims to generate a video of a certain reference identity following a particular motion sequence. Existing animation works typically employ the frame-warping technique to animate the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zhongcong Xu , Jianfeng Zhang , Jun Hao Liew , Hanshu Yan , Jia-Wei Liu , Chenxu Zhang , Jiashi Feng , Mike Zheng Shou

Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a real-time framework for infinite, 3D-consistent video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ke Zhang , Yiqun Mei , Jiacong Xu , Vishal M. Patel

Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics driven by actions, with visual observations providing partial information about the state. Recent video world models attempt to learn this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zhen Li , Zian Meng , Shuwei Shi , Wenshuo Peng , Yuwei Wu , Bo Zheng , Chuanhao Li , Kaipeng Zhang

This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Wenqiang Sun , Haiyu Zhang , Haoyuan Wang , Junta Wu , Zehan Wang , Zhenwei Wang , Yunhong Wang , Jun Zhang , Tengfei Wang , Chunchao Guo

With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal…

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects in isolation, as…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yicong Hong , Yiqun Mei , Chongjian Ge , Yiran Xu , Yang Zhou , Sai Bi , Yannick Hold-Geoffroy , Mike Roberts , Matthew Fisher , Eli Shechtman , Kalyan Sunkavalli , Feng Liu , Zhengqi Li , Hao Tan

Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle with a lack of spatial persistence and insufficient visual…

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Taiye Chen , Xun Hu , Zihan Ding , Chi Jin

We propose Infinite-World, a robust interactive world model capable of maintaining coherent visual memory over 1000+ frames in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Ruiqi Wu , Xuanhua He , Meng Cheng , Tianyu Yang , Yong Zhang , Zhuoliang Kang , Xunliang Cai , Xiaoming Wei , Chunle Guo , Chongyi Li , Ming-Ming Cheng

Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Quanhao Li , Zhen Xing , Rui Wang , Hui Zhang , Qi Dai , Zuxuan Wu

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

‹ Prev 1 2 3 10 Next ›