中文
相关论文

相关论文: INSPATIO-WORLD: A Real-Time 4D World Simulator via…

200 篇论文

We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential frame generation and incur substantial latency due to window-level processing,…

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu , Yanyong Zhang

We introduce InfinityStar, a unified spacetime autoregressive framework for high-resolution image and dynamic video synthesis. Building on the recent success of autoregressive modeling in both vision and language, our purely discrete…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Jinlai Liu , Jian Han , Bin Yan , Hui Wu , Fengda Zhu , Xing Wang , Yi Jiang , Bingyue Peng , Zehuan Yuan

Generating high-quality 4D objects with spatial-temporal consistency is still formidable. Existing diffusion-based methods often struggle with spatial-temporal inconsistency, as they fail to leverage outputs from all previous timesteps to…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Liying Yang , Jialun Liu , Jiakui Hu , Chenhao Guan , Haibin Huang , Fangqiu Yi , Chi Zhang , Yanyan Liang

Autonomous driving systems struggle with complex scenarios due to limited access to diverse, extensive, and out-of-distribution driving data which are critical for safe navigation. World models offer a promising solution to this challenge;…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Xi Guo , Chenjing Ding , Haoxuan Dou , Xin Zhang , Weixuan Tang , Wei Wu

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Guangyuan Li , Bo Li , Jinwei Chen , Xiaobin Hu , Lei Zhao , Peng-Tao Jiang

We propose Infinite-World, a robust interactive world model capable of maintaining coherent visual memory over 1000+ frames in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Ruiqi Wu , Xuanhua He , Meng Cheng , Tianyu Yang , Yong Zhang , Zhuoliang Kang , Xunliang Cai , Xiaoming Wei , Chunle Guo , Chongyi Li , Ming-Ming Cheng

High-quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high-quality dynamic scenes typically require dense…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Weihong Pan , Xiaoyu Zhang , Zhuang Zhang , Zhichao Ye , Nan Wang , Haomin Liu , Guofeng Zhang

Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D representation of the environment, meaning 3D consistency must…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Samuel Garcin , Thomas Walker , Steven McDonagh , Tim Pearce , Hakan Bilen , Tianyu He , Kaixin Wang , Jiang Bian

We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Hyeongju Mun , In-Hwan Jin , Sohyeong Kim , Kyeongbo Kong

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically, existing 4D generative models directly embed macro scale…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Haonan Wang , Hanyu Zhou , Tao Gu , Luxin Yan

Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Xiaotao Hu , Wei Yin , Mingkai Jia , Junyuan Deng , Xiaoyang Guo , Qian Zhang , Xiaoxiao Long , Ping Tan

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Ying Yang , Zhengyao Lv , Tianlin Pan , Haofan Wang , Binxin Yang , Hubery Yin , Chen Li , Ziwei Liu , Chenyang Si

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal context window sizes, these models often struggle to…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Tong Wu , Shuai Yang , Ryan Po , Yinghao Xu , Ziwei Liu , Dahua Lin , Gordon Wetzstein

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Rui Xie , Yinhong Liu , Penghao Zhou , Chen Zhao , Jun Zhou , Kai Zhang , Zhenyu Zhang , Jian Yang , Zhenheng Yang , Ying Tai

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yisu Zhang , Chenjie Cao , Tengfei Wang , Xuhui Zuo , Junta Wu , Jianke Zhu , Chunchao Guo

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Sen Wang , Jingyi Tian , Le Wang , Zhimin Liao , Jiayi Li , Huaiyi Dong , Kun Xia , Sanping Zhou , Wei Tang , Hua Gang

With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal…

Semantic occupancy has emerged as a powerful representation in world models for its ability to capture rich spatial semantics. However, most existing occupancy world models rely on static and fixed embeddings or grids, which inherently…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Chenxu Dang , Haiyan Liu , Jason Bao , Pei An , Xinyue Tang , PanAn , Jie Ma , Bingchuan Sun , Yan Wang
‹ 上一页 1 2 3 10 下一页 ›