English
Related papers

Related papers: Towards World Simulator: Crafting Physical Commons…

200 papers

While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core challenge. In this work, we systematically study Newtonian…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Xiangyu Bai , He Liang , Bishoy Galoaa , Utsav Nandi , Shayda Moezzi , Yuhang He , Sarah Ostadabbas

Driven by the growing capacity and training scale, Text-to-Video (T2V) generation models have recently achieved substantial progress in video quality, length, and instruction-following capability. However, whether these models can…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zeqing Wang , Keze Wang , Lei Zhang

Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popularity of the technology. This progress has created a growing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Zelu Qi , Ping Shi , Shuqi Wang , Chaoyang Zhang , Fei Zhao , Zefeng Ying , Da Pan , Xi Yang , Zheqi He , Teng Dai

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

Video generation models are increasingly used as world simulators for storytelling, simulation, and embodied AI. As these models advance, a key question arises: do generated videos obey the physical laws of the real world? Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Qin Zhang , Peiyu Jing , Hong-Xing Yu , Fangqiang Ding , Fan Nie , Weimin Wang , Yilun Du , James Zou , Jiajun Wu , Bing Shuai

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Keming Wu , Yijing Cui , Wenhan Xue , Qijie Wang , Xuan Luo , Zhiyuan Feng , Zuhao Yang , Sudong Wang , Sicong Jiang , Haowei Zhu , Zihan Wang , Ping Nie , Wenhu Chen , Bin Wang

The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making the existing benchmarks increasingly insufficient to evaluate state-of-the-art T2V models.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Zeqing Wang , Xinyu Wei , Bairui Li , Zhen Guo , Jinrui Zhang , Hongyang Wei , Keze Wang , Lei Zhang

Recent text-to-video generation models have made remarkable progress in visual realism, motion fidelity, and text-video alignment, yet they still struggle to produce socially coherent behavior. Unlike humans, who readily infer intentions,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Wenshuo Peng , Gongxuan Wang , Tianmeng Yang , Chuanhao Li , Xiaojie Xu , Hui He , Kaipeng Zhang

Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and…

Artificial Intelligence · Computer Science 2026-02-12 Lintao Wang , Encheng Su , Jiaqi Liu , Pengze Li , Jiabei Xiao , Wenlong Zhang , Xinnan Dai , Xi Chen , Yuan Meng , Lei Bai , Wanli Ouyang , Shixiang Tang , Aoran Wang , Xinzhu Ma

Recent progress in text-to-image (T2I) generation underscores the importance of reliable benchmarks in evaluating how accurately generated images reflect the semantics of their textual prompt. However, (1) existing benchmarks lack the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Yibin Wang , Zhimin Li , Yuhang Zang , Jiazi Bu , Yujie Zhou , Yi Xin , Junjun He , Chunyu Wang , Qinglin Lu , Cheng Jin , Jiaqi Wang

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Ziwei Zhou , Zeyuan Lai , Rui Wang , Yifan Yang , Zhen Xing , Yuqing Yang , Qi Dai , Lili Qiu , Chong Luo

We present PhysGen, a novel image-to-video generation method that converts a single image and an input condition (e.g., force and torque applied to an object in the image) to produce a realistic, physically plausible, and temporally…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Shaowei Liu , Zhongzheng Ren , Saurabh Gupta , Shenlong Wang

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals are domain-specific…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Saurabh Pathak , Elahe Arani , Mykola Pechenizkiy , Bahram Zonooz

Video generation models have shown strong potential as world models for autonomous driving simulation. However, existing approaches are primarily trained on real-world driving datasets, which mostly contain natural and safe driving…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jiawei Zhou , Zhenxin Zhu , Lingyi Du , Linye Lyu , Lijun Zhou , Zhanqian Wu , Hongcheng Luo , Zhuotao Tian , Bing Wang , Guang Chen , Hangjun Ye , Haiyang Sun , Yu Li

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Harold Haodong Chen , Disen Lan , Wen-Jie Shu , Qingyang Liu , Zihan Wang , Sirui Chen , Wenkai Cheng , Kanghao Chen , Hongfei Zhang , Zixin Zhang , Rongjin Guo , Yu Cheng , Ying-Cong Chen

Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols either emphasize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Tianyang Han , Junhao Su , Junjie Hu , Peizhen Yang , Hengyu Shi , Junfeng Luo , Jialin Gao

Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding temporal decay and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Kaijie Chen , Zihao Lin , Zhiyang Xu , Ying Shen , Yuguang Yao , Joy Rimchala , Jiaxin Zhang , Lifu Huang

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Physics-aware symbolic simulation of 3D scenes is critical for robotics, embodied AI, and scientific computing, requiring models to understand natural language descriptions of physical phenomena and translate them into executable simulation…

Robotics · Computer Science 2026-04-28 Tianyidan Xie , Peiyu Wang , Yuyi Qian , Yuxuan Wang , Rui Ma , Ying Tai , Song Wu , Qian Wang , Lanjun Wang , Zili Yi

Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models support these abilities is insufficiently understood. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Fengzhe Zhou , Jiannan Huang , Jialuo Li , Deva Ramanan , Humphrey Shi