English
Related papers

Related papers: How Far is Video Generation from World Model: A Ph…

200 papers

Large-scale pre-trained video generation models excel in content creation but are not reliable as physically accurate world simulators out of the box. This work studies the process of post-training these models for accurate world modeling…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Chenyu Li , Oscar Michel , Xichen Pan , Sainan Liu , Mike Roberts , Saining Xie

The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena. Despite its growing popularity, there is a lack of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Xuanyi Li , Daquan Zhou , Chenxu Zhang , Shaodong Wei , Qibin Hou , Ming-Ming Cheng

General world models represent a crucial pathway toward achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications ranging from virtual environments to decision-making systems. Recently, the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Zheng Zhu , Xiaofeng Wang , Wangbo Zhao , Chen Min , Bohan Li , Nianchen Deng , Min Dou , Yuqi Wang , Botian Shi , Kai Wang , Chi Zhang , Yang You , Zhaoxiang Zhang , Dawei Zhao , Liang Xiao , Jian Zhao , Jiwen Lu , Guan Huang

Sora is a text-to-video generative AI model, released by OpenAI in February 2024. The model is trained to generate videos of realistic or imaginative scenes from text instructions and show potential in simulating the physical world. Based…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Yixin Liu , Kai Zhang , Yuan Li , Zhiling Yan , Chujie Gao , Ruoxi Chen , Zhengqing Yuan , Yue Huang , Hanchi Sun , Jianfeng Gao , Lifang He , Lichao Sun

With impressive achievements made, artificial intelligence is on the path forward to artificial general intelligence. Sora, developed by OpenAI, which is capable of minute-level world-simulative abilities can be considered as a milestone on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Rui Sun , Yumin Zhang , Tejal Shah , Jiahao Sun , Shuoying Zhang , Wenqi Li , Haoran Duan , Bo Wei , Rajiv Ranjan

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Faraz Waseem , Muhammad Shahzad

AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics -- or, alternatively,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Saman Motamed , Laura Culp , Kevin Swersky , Priyank Jaini , Robert Geirhos

Video generators are increasingly evaluated as potential world models, which requires them to encode and understand physical laws. We investigate their representation of a fundamental law: gravity. Out-of-the-box video generators…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Varun Varma Thozhiyoor , Shivam Tripathi , Venkatesh Babu Radhakrishnan , Anand Bhattad

The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential…

Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents and their environments conditioned on multi-modal user…

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

Artificial Intelligence · Computer Science 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao

We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the ability to observe other parameters such as the object velocities or the forces applied to them. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Gabrijel Boduljak , Laurynas Karazija , Iro Laina , Christian Rupprecht , Andrea Vedaldi

Recent advancements in video generation have witnessed significant progress, especially with the rapid advancement of diffusion models. Despite this, their deficiencies in physical cognition have gradually received widespread attention -…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Minghui Lin , Xiang Wang , Yishan Wang , Shu Wang , Fengqi Dai , Pengxiang Ding , Cunxiang Wang , Zhengrong Zuo , Nong Sang , Siteng Huang , Donglin Wang

Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is…

Robotics · Computer Science 2025-08-04 Junbang Liang , Pavel Tokmakov , Ruoshi Liu , Sruthi Sudhakar , Paarth Shah , Rares Ambrus , Carl Vondrick

The advent of text-to-video generation models has revolutionized content creation as it produces high-quality videos from textual prompts. However, concerns regarding inherent biases in such models have prompted scrutiny, particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Mohammad Nadeem , Shahab Saquib Sohail , Erik Cambria , Björn W. Schuller , Amir Hussain

Animation has gained significant interest in the recent film and TV industry. Despite the success of advanced video generation models like Sora, Kling, and CogVideoX in generating natural videos, they lack the same effectiveness in handling…

Video generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Dacheng Li , Yunhao Fang , Yukang Chen , Shuo Yang , Shiyi Cao , Justin Wong , Michael Luo , Xiaolong Wang , Hongxu Yin , Joseph E. Gonzalez , Ion Stoica , Song Han , Yao Lu

A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicitly represent such…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Sonia Joseph , Quentin Garrido , Randall Balestriero , Matthew Kowal , Thomas Fel , Shahab Bakhtiari , Blake Richards , Mike Rabbat

Video Diffusion Models (VDMs) have emerged as powerful generative tools, capable of synthesizing high-quality spatiotemporal content. Yet, their potential goes far beyond mere video generation. We argue that the training dynamics of VDMs,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro
‹ Prev 1 2 3 10 Next ›