中文
相关论文

相关论文: Improving the Physics of Video Generation with VJE…

200 篇论文

State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we…

Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To this end, we propose an iterative self-refinement framework…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yang Liu , Xilin Zhao , Peisong Wen , Siran Dai , Qingming Huang

Current video generation models produce high-quality aesthetic videos but often struggle to learn representations of real-world physics dynamics, resulting in artifacts such as unnatural object collisions, inconsistent gravity, and temporal…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Siddarth Nilol Kundur Satish , Devesh Jaiswal , Hongyu Chen , Abhishek Bakshi

AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics -- or, alternatively,…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Saman Motamed , Laura Culp , Kevin Swersky , Priyank Jaini , Robert Geirhos

Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Xiangdong Zhang , Jiaqi Liao , Shaofeng Zhang , Fanqing Meng , Xiangpeng Wan , Junchi Yan , Yu Cheng

Generative AI models, particularly Text-to-Video (T2V) systems, offer a promising avenue for transforming science education by automating the creation of engaging and intuitive visual explanations. In this work, we take a first step toward…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Megha Mariam K. M , Aditya Arun , Zakaria Laskar , C. V. Jawahar

Recent advances in diffusion-based and autoregressive video generation models have achieved remarkable visual realism. However, these models typically lack accurate physical alignment, failing to replicate real-world dynamics in object…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Tao Feng , Xianbing Zhao , Zhenhua Chen , Tien Tsin Wong , Hamid Rezatofighi , Gholamreza Haffari , Lizhen Qu

Current image generation models produce visually compelling but scientifically implausible images, exposing a fundamental gap between visual fidelity and physical realism. In this work, we introduce ScienceT2I, an expert-annotated dataset…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Jialuo Li , Wenhao Chai , Xingyu Fu , Haiyang Xu , Saining Xie

Driven by the growing capacity and training scale, Text-to-Video (T2V) generation models have recently achieved substantial progress in video quality, length, and instruction-following capability. However, whether these models can…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Zeqing Wang , Keze Wang , Lei Zhang

Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult to scale, and still prone to producing implausible motions…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Yutong Hao , Chen Chen , Ajmal Saeed Mian , Chang Xu , Daochang Liu

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities, the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous…

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Hritik Bansal , Clark Peng , Yonatan Bitton , Roman Goldenberg , Aditya Grover , Kai-Wei Chang

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Xindi Yang , Baolu Li , Yiming Zhang , Zhenfei Yin , Lei Bai , Liqian Ma , Zhiyong Wang , Jianfei Cai , Tien-Tsin Wong , Huchuan Lu , Xu Jia

While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Mingxin Liu , Shuran Ma , Shibei Meng , Xiangyu Zhao , Zicheng Zhang , Shaofeng Zhang , Zhihang Zhong , Peixian Chen , Haoyu Cao , Xing Sun , Haodong Duan , Xue Yang

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Lanxiang Hu , Abhilash Shankarampeta , Yixin Huang , Zilin Dai , Haoyang Yu , Yujie Zhao , Haoqiang Kang , Daniel Zhao , Tajana Rosing , Hao Zhang

Recent advancements in video generation have witnessed significant progress, especially with the rapid advancement of diffusion models. Despite this, their deficiencies in physical cognition have gradually received widespread attention -…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Minghui Lin , Xiang Wang , Yishan Wang , Shu Wang , Fengqi Dai , Pengxiang Ding , Cunxiang Wang , Zhengrong Zuo , Nong Sang , Siteng Huang , Donglin Wang

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art creation and user…

机器学习 · 计算机科学 2025-05-02 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle to grasp abstract physical principles and generate videos…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jing Wang , Ao Ma , Ke Cao , Jun Zheng , Zhanjie Zhang , Jiasong Feng , Shanyuan Liu , Yuhang Ma , Bo Cheng , Dawei Leng , Yuhui Yin , Xiaodan Liang

Recent advances in text-to-video (T2V) generation have achieved good visual quality, yet synthesizing videos that faithfully follow physical laws remains an open challenge. Existing methods mainly based on graphics or prompt extension…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Yuanhao Cai , Kunpeng Li , Menglin Jia , Jialiang Wang , Junzhe Sun , Feng Liang , Weifeng Chen , Felix Juefei-Xu , Chu Wang , Ali Thabet , Xiaoliang Dai , Xuan Ju , Alan Yuille , Ji Hou

State-of-the-art text-to-video (T2V) generators frequently violate physical laws despite high visual quality. We show this stems from insufficient physical constraints in prompts rather than model limitations: manually adding physics…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Shang Wu , Chenwei Xu , Zhuofan Xia , Weijian Li , Lie Lu , Pranav Maneriker , Fan Du , Manling Li , Han Liu
‹ 上一页 1 2 3 10 下一页 ›