English
Related papers

Related papers: KeyWorld: Key Frame Reasoning Enables Effective an…

200 papers

This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward…

World foundation models aim to simulate the evolution of the real world with physically plausible behavior. Unlike prior methods that handle spatial and temporal correlations separately, we propose RAYNOVA, a geometry-agonistic multiview…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Yichen Xie , Chensheng Peng , Mazen Abdelfattah , Yihan Hu , Jiezhi Yang , Eric Higgins , Ryan Brigden , Masayoshi Tomizuka , Wei Zhan

World Foundation Models (WFMs) offer remarkable visual dynamics simulation capabilities, yet their application to precise robotic control remains limited by the gap between generative realism and control-oriented precision. While existing…

Robotics · Computer Science 2025-12-04 Yuhang Huang , Shilong Zou , Jiazhao Zhang , Xinwang Liu , Ruizhen Hu , Kai Xu

The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs…

Robotics · Computer Science 2026-03-31 Andrew Choi , Xinjie Wang , Zhizhong Su , Wei Xu

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Weijie Wang , Xiaoxuan He , Youping Gu , Yifan Yang , Zeyu Zhang , Yefei He , Yanbo Ding , Xirui Hu , Donny Y. Chen , Zhiyuan He , Yuqing Yang , Bohan Zhuang

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Nate Gillman , Yinghua Zhou , Zitian Tang , Evan Luo , Arjan Chakravarthy , Daksh Aggarwal , Michael Freeman , Charles Herrmann , Chen Sun

State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we…

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yubin Chen , Xuyang Guo , Zhenmei Shi , Zhao Song , Jiahao Zhang

Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks rely on recognition-style protocols such as Visual Question Answering (VQA) and Violation of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Jiarong Liang , Max Ku , Ka-Hei Hui , Ping Nie , Wenhu Chen

Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectively capture the temporal dependencies and redundancies…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Xiang Fan , Xiaohang Sun , Kushan Thakkar , Zhu Liu , Vimal Bhat , Ranjay Krishna , Xiang Hao

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the…

Robotics · Computer Science 2025-07-30 Fangqi Zhu , Hongtao Wu , Song Guo , Yuxiao Liu , Chilam Cheang , Tao Kong

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xingtai Gui , Meijie Zhang , Tianyi Yan , Wencheng Han , Jiahao Gong , Feiyang Tan , Cheng-zhong Xu , Jianbing Shen

Deep learning has rapidly transformed the state of the art algorithms used to address a variety of problems in computer vision and robotics. These breakthroughs have relied upon massive amounts of human annotated training data. This time…

Computer Vision and Pattern Recognition · Computer Science 2017-02-28 Matthew Johnson-Roberson , Charles Barto , Rounak Mehta , Sharath Nittur Sridhar , Karl Rosaen , Ram Vasudevan

World modelling, i.e. building a representation of the rules that govern the world so as to predict its evolution, is an essential ability for any agent interacting with the physical world. Recent applications of the Transformer…

Machine Learning · Computer Science 2024-05-31 Francesco Petri , Luigi Asprino , Aldo Gangemi

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

Robotics · Computer Science 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen

World models are central to building agents that can reason, plan, and generalize beyond their training data. However, research on world models is currently fragmented, with disparate codebases, data pipelines, and evaluation protocols…

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies…

Robotics · Computer Science 2026-01-08 Wenlong Huang , Yu-Wei Chao , Arsalan Mousavian , Ming-Yu Liu , Dieter Fox , Kaichun Mo , Li Fei-Fei

The concept of world models has garnered significant attention due to advancements in multimodal large language models such as GPT-4 and video generation models such as Sora, which are central to the pursuit of artificial general…

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Ying Yang , Zhengyao Lv , Tianlin Pan , Haofan Wang , Binxin Yang , Hubery Yin , Chen Li , Ziwei Liu , Chenyang Si