English
Related papers

Related papers: Inference-time Physics Alignment of Video Generati…

200 papers

Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jie Liu , Gongye Liu , Jiajun Liang , Ziyang Yuan , Xiaokun Liu , Mingwu Zheng , Xiele Wu , Qiulin Wang , Menghan Xia , Xintao Wang , Xiaohong Liu , Fei Yang , Pengfei Wan , Di Zhang , Kun Gai , Yujiu Yang , Wanli Ouyang

Action-conditioned robot world models generate future video frames of the manipulated scene given a robot action sequence, offering a promising alternative for simulating tasks that are difficult to model with traditional physics engines.…

Robotics · Computer Science 2026-03-27 Jai Bardhan , Patrik Drozdik , Josef Sivic , Vladimir Petrik

Learning predictive models from high-dimensional sensory observations is fundamental for cyber-physical systems, yet the latent representations learned by standard world models lack physical interpretability. This limits their reliability,…

Machine Learning · Computer Science 2026-04-07 Zhenjiang Mao , Mrinall Eashaan Umasudhan , Ivan Ruchkin

Large-scale pre-trained video generation models excel in content creation but are not reliable as physically accurate world simulators out of the box. This work studies the process of post-training these models for accurate world modeling…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Chenyu Li , Oscar Michel , Xichen Pan , Sainan Liu , Mike Roberts , Saining Xie

Recent advances in video generation have enabled the synthesis of videos with strong temporal consistency and impressive visual quality, marking a crucial step toward vision foundation models. To evaluate these video generation models,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Xuming He , Zehao Fan , Hengjia Li , Fan Zhuo , Hankun Xu , Senlin Cheng , Di Weng , Haifeng Liu , Can Ye , Boxi Wu

Vision-Language-Action (VLA) models trained via imitation learning suffer from significant performance degradation in data-scarce scenarios due to their reliance on large-scale demonstration datasets. Although reinforcement learning…

Robotics · Computer Science 2026-04-28 Junjin Xiao , Yandan Yang , Xinyuan Chang , Ronghan Chen , Feng Xiong , Mu Xu , Wei-Shi Zheng , Qing Zhang

Computer vision is hard because of a large variability in lighting, shape, and texture; in addition the image signal is non-additive due to occlusion. Generative models promised to account for this variability by accurately modelling the…

Computer Vision and Pattern Recognition · Computer Science 2015-03-10 Varun Jampani , Sebastian Nowozin , Matthew Loper , Peter V. Gehler

Recent advances in text-to-video generation, particularly with autoregressive models, have enabled the synthesis of high-quality videos depicting individual scenes. However, extending these models to generate long, cross-scene videos…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Xueji Fang , Liyuan Ma , Zhiyang Chen , Mingyuan Zhou , Guo-jun Qi

Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Designing such reward functions can be challenging and may not…

Machine Learning · Computer Science 2026-04-06 Qi Wang , Mian Wu , Yuyang Zhang , Mingqi Yuan , Wenyao Zhang , Haoxiang You , Yunbo Wang , Xin Jin , Xiaokang Yang , Wenjun Zeng

Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of a pretrained Diffusion Transformer (DiT) and find that physically plausible and implausible videos are…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Chujun Tang , Lei Zhong , Fangqiang Ding

Reinforcement Learning with Verifiable Rewards (RLVR) has substantially advanced the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, the rapid progress of MLLMs is outpacing the complexity of existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Zefeng He , Xiaoye Qu , Yafu Li , Siyuan Huang , Daizong Liu , Yu Cheng

Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Meiqi Wu , Bingze Song , Ruimin Lin , Chen Zhu , Xiaokun Feng , Jiahong Wu , Xiangxiang Chu , Kaiqi Huang

The "Reason-Then-Respond" paradigm, enhanced by Reinforcement Learning, has shown great promise in advancing Multimodal Large Language Models. However, its application to the video domain has led to specialized models that excel at either…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Xinlong Chen , Yuanxing Zhang , Yushuo Guan , Weihong Lin , Zekun Wang , Bohan Zeng , Yang Shi , Sihan Yang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

Long video generation remains a challenging and compelling topic in computer vision. Diffusion based models, among the various approaches to video generation, have achieved state of the art quality with their iterative denoising procedures.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Siyang Zhang , Harry Yang , Ser-Nam Lim

We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one. Our approach models the generation of safe responses as a constrained Markov…

Machine Learning · Computer Science 2025-06-23 Xiaotong Ji , Shyam Sundhar Ramesh , Matthieu Zimmer , Ilija Bogunovic , Jun Wang , Haitham Bou Ammar

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diffusion models.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zixuan Wang , Yixin Hu , Haolan Wang , Feng Chen , Yan Liu , Wen Li , Yinjie Lei

Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first framework to scale…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Shubo Lin , Xuanyang Zhang , Wei Cheng , Weiming Hu , Gang Yu , Jin Gao

Flow matching is a recent framework to train generative models that exhibits impressive empirical performance while being relatively easier to train compared with diffusion-based models. Despite its advantageous properties, prior methods…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Quan Dao , Hao Phung , Binh Nguyen , Anh Tran

In this paper, we propose a novel end-to-end architecture that could generate a variety of plausible video sequences correlating two given discontinuous frames. Our work is inspired by the human ability of inference. Specifically, given two…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Weimian Li , Baoyang Chen , Wenmin Wang

Advancements in language foundation models have primarily fueled the recent surge in artificial intelligence. In contrast, generative learning of non-textual modalities, especially videos, significantly trails behind language modeling. This…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Lijun Yu