English
Related papers

Related papers: IC-World: In-Context Generation for Shared World M…

200 papers

World models serve as essential building blocks toward Artificial General Intelligence (AGI), enabling intelligent agents to predict future states and plan actions by simulating complex physical interactions. However, existing interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junyi Chen , Haoyi Zhu , Xianglong He , Yifan Wang , Jianjun Zhou , Wenzheng Chang , Yang Zhou , Zizun Li , Zhoujie Fu , Jiangmiao Pang , Tong He

Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines are based on implicit generative models, which are hard to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Wenhao Hu , Xuexiang Wen , Xi Li , Gaoang Wang

In neural video codecs, current state-of-the-art methods typically adopt multi-scale motion compensation to handle diverse motions. These methods estimate and compress either optical flow or deformable offsets to reduce inter-frame…

Multimedia · Computer Science 2024-12-03 Yongqi Zhai , Jiayu Yang , Wei Jiang , Chunhui Yang , Luyang Tang , Ronggang Wang

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yang Zhou , Hao Shao , Letian Wang , Zhuofan Zong , Hongsheng Li , Steven L. Waslander

Adapting pretrained video generation models into controllable world models via latent actions is a promising step towards creating generalist world models. The dominant paradigm adopts a two-stage approach that trains latent action model…

Machine Learning · Computer Science 2026-04-07 Yucen Wang , Fengming Zhang , De-Chuan Zhan , Li Zhao , Kaixin Wang , Jiang Bian

Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateless" video architectures and classic state-centric world model…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Luozhou Wang , Zhifei Chen , Yihua Du , Dongyu Yan , Wenhang Ge , Guibao Shen , Xinli Xu , Leyi Wu , Man Chen , Tianshuo Xu , Peiran Ren , Xin Tao , Pengfei Wan , Ying-Cong Chen

What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificial environments by imagining all content. We present Seoul…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Junyoung Seo , Hyunwook Choi , Minkyung Kwon , Jinhyeok Choi , Siyoon Jin , Gayoung Lee , Junho Kim , JoungBin Lee , Geonmo Gu , Dongyoon Han , Sangdoo Yun , Seungryong Kim , Jin-Hwa Kim

The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation methods conventionally perform in RGB pixel space, with…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Cuifeng Shen , Yulu Gan , Chen Chen , Xiongwei Zhu , Lele Cheng , Tingting Gao , Jinzhi Wang

World models are a powerful paradigm in AI and robotics, enabling agents to reason about the future by predicting visual observations or compact latent states. The 1X World Model Challenge introduces an open-source benchmark of real-world…

Machine Learning · Computer Science 2025-10-09 Riccardo Mereu , Aidan Scannell , Yuxin Hou , Yi Zhao , Aditya Jitta , Antonio Dominguez , Luigi Acerbi , Amos Storkey , Paul Chang

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xingtai Gui , Meijie Zhang , Tianyi Yan , Wencheng Han , Jiahao Gong , Feiyang Tan , Cheng-zhong Xu , Jianbing Shen

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Qihang Zhang , Shuangfei Zhai , Miguel Angel Bautista , Kevin Miao , Alexander Toshev , Joshua Susskind , Jiatao Gu

The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these…

Machine Learning · Computer Science 2020-10-06 Zhixuan Lin , Yi-Fu Wu , Skand Peri , Bofeng Fu , Jindong Jiang , Sungjin Ahn

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yisu Zhang , Chenjie Cao , Tengfei Wang , Xuhui Zuo , Junta Wu , Jianke Zhu , Chunchao Guo

In this work, we focus on a challenging task: synthesizing multiple imaginary videos given a single image. Major problems come from high dimensionality of pixel space and the ambiguity of potential motions. To overcome those problems, we…

Computer Vision and Pattern Recognition · Computer Science 2017-06-16 Baoyang Chen , Wenmin Wang , Jinzhuo Wang , Xiongtao Chen

Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitation, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Tengjiao Yin , Jinglei Shi , Heng Guo , Xi Wang

Building on the momentum of image generation diffusion models, there is an increasing interest in video-based diffusion models. However, video generation poses greater challenges due to its higher-dimensional nature, the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Aimon Rahman , Malsha V. Perera , Vishal M. Patel

Image-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Jie Tian , Xiaoye Qu , Zhenyi Lu , Wei Wei , Sichen Liu , Yu Cheng

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video generation with evolving…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Jiahao Wang , Luoxin Ye , TaiMing Lu , Junfei Xiao , Jiahan Zhang , Yuxiang Guo , Xijun Liu , Rama Chellappa , Cheng Peng , Alan Yuille , Jieneng Chen

Generative inbetweening aims to generate intermediate frame sequences by utilizing two key frames as input. Although remarkable progress has been made in video generation models, generative inbetweening still faces challenges in maintaining…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Tianyi Zhu , Dongwei Ren , Qilong Wang , Xiaohe Wu , Wangmeng Zuo