English
Related papers

Related papers: WoW: Towards a World omniscient World model Throug…

200 papers

Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant…

Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requires extensive modeling and manual setup, and more…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Quankai Gao , Jiawei Yang , Qiangeng Xu , Le Chen , Yue Wang

The predictive functions that permit humans to infer their body state by sensorimotor integration are critical to perform safe interaction in complex environments. These functions are adaptive and robust to non-linear actuators and noisy…

Robotics · Computer Science 2019-10-24 Pablo Lanillos , Gordon Cheng

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Fuxiang Yang , Donglin Di , Lulu Tang , Xuancheng Zhang , Lei Fan , Hao Li , Chen Wei , Tonghua Su , Baorui Ma

World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are often either learned directly from pixels with limited…

Artificial Intelligence · Computer Science 2026-05-26 Minghao Fu , Fan Feng , Nicklas Hansen , Biwei Huang

Agents that understand objects and their interactions can learn policies that are more robust and transferable. However, most object-centric RL methods factor state by individual objects while leaving interactions implicit. We introduce the…

Machine Learning · Computer Science 2025-11-05 Fan Feng , Phillip Lippe , Sara Magliacane

OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. However, the ability of video generation models to discover such laws purely from visual data without human…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Bingyi Kang , Yang Yue , Rui Lu , Zhijie Lin , Yang Zhao , Kaixin Wang , Gao Huang , Jiashi Feng

World models have been developed to support sample-efficient deep reinforcement learning agents. However, it remains challenging for world models to accurately replicate environments that are high-dimensional, non-stationary, and composed…

Machine Learning · Computer Science 2026-03-31 Yosuke Nishimoto , Takashi Matsubara

Humans naturally build mental models of object interactions and dynamics, allowing them to imagine how their surroundings will change if they take a certain action. While generative models today have shown impressive results on…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Sruthi Sudhakar , Ruoshi Liu , Basile Van Hoorick , Carl Vondrick , Richard Zemel

Robotic world models are a promising paradigm for forecasting future environment states, yet their inference speed and the physical plausibility of generated trajectories remain critical bottlenecks, limiting their real-world applications.…

Robotics · Computer Science 2025-09-26 Sibo Li , Qianyue Hao , Yu Shang , Yong Li

A robot's instantaneous sensory observations do not always reveal task-relevant state information. Under such partial observability, optimal behavior typically involves explicitly acting to gain the missing information. Today's standard…

Humans understand the world through the integration of multiple sensory modalities, enabling them to perceive, reason about, and imagine dynamic physical processes. Inspired by this capability, multimodal foundation models (MFMs) have…

Artificial Intelligence · Computer Science 2025-10-07 Xuehai He

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Yiran Qin , Zhelun Shi , Jiwen Yu , Xijun Wang , Enshen Zhou , Lijun Li , Zhenfei Yin , Xihui Liu , Lu Sheng , Jing Shao , Lei Bai , Wanli Ouyang , Ruimao Zhang

From just a glance, humans can make rich predictions about the future state of a wide range of physical systems. On the other hand, modern approaches from engineering, robotics, and graphics are often restricted to narrow domains and…

Computer Vision and Pattern Recognition · Computer Science 2017-06-06 Nicholas Watters , Andrea Tacchetti , Theophane Weber , Razvan Pascanu , Peter Battaglia , Daniel Zoran

Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame world modeling as novel-view synthesis or future-frame prediction, emphasizing visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yifan Yin , Zehao Wen , Jieneng Chen , Zehan Zheng , Nanru Dai , Haojun Shi , Suyu Ye , Aydan Huang , Zheyuan Zhang , Alan Yuille , Jianwen Xie , Ayush Tewari , Tianmin Shu

Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically…

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu

Learning predictive models from high-dimensional sensory observations is fundamental for cyber-physical systems, yet the latent representations learned by standard world models lack physical interpretability. This limits their reliability,…

Machine Learning · Computer Science 2026-04-07 Zhenjiang Mao , Mrinall Eashaan Umasudhan , Ivan Ruchkin

We present a novel study on enhancing the capability of preserving the content in world models, focusing on a property we term World Stability. Recent diffusion-based generative models have advanced the synthesis of immersive and realistic…

Machine Learning · Computer Science 2025-03-12 Soonwoo Kwon , Jin-Young Kim , Hyojun Go , Kyungjune Baek

Autonomous robots need to be able to adapt to unforeseen situations and to acquire new skills through trial and error. Reinforcement learning in principle offers a suitable methodological framework for this kind of autonomous learning.…

Robotics · Computer Science 2016-08-02 Nikolas J. Hemion
‹ Prev 1 4 5 6 7 8 10 Next ›