English
Related papers

Related papers: MineWorld: a Real-Time and Open-Source Interactive…

200 papers

Humanoid robots, with their human-like form, are uniquely suited for interacting in environments built for people. However, enabling humanoids to reason, plan, and act in complex open-world settings remains a challenge. World models, models…

Robotics · Computer Science 2025-07-10 Muhammad Qasim Ali , Aditya Sridhar , Shahbuland Matiana , Alex Wong , Mohammad Al-Sharman

World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become unreliable once the robot encounters states outside the training distribution, limiting…

Robotics · Computer Science 2026-05-18 Tuo An , Jindou Jia , Gen Li , Jingliang Li , Chuhao Zhou , Pengfei Liu , Bofan Lyu , Jiaqi Bai , Xinying Guo , Geng Li , Jianfei Yang

World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Zeqi Xiao , Yushi Lan , Yifan Zhou , Wenqi Ouyang , Shuai Yang , Yanhong Zeng , Xingang Pan

Social navigation requires robots to act safely in dynamic human environments. Effective behavior demands thinking ahead: reasoning about how the scene and pedestrians evolve under different robot actions rather than reacting to current…

Robotics · Computer Science 2026-03-19 Tianshuai Hu , Zeying Gong , Lingdong Kong , XiaoDong Mei , Yiyi Ding , Qi Zeng , Ao Liang , Rong Li , Yangyi Zhong , Junwei Liang

Recent video-based world models have made pixel-space environments interactive at the camera level: users can navigate viewpoints while the model generates coherent visual continuations. Yet their action spaces remain incomplete: users can…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Bohai Gu , Taiyi Wu , Yueyang Yuan , Jian Liu , Xiaocheng Lu , Dazhao Du , Jie Zhang , Jinxiang Lai , Shuai Yang , Xiaotong Zhao , Alan Zhao , Song Guo

Supporting real-time interactions between human controllers and remote devices remains a challenging goal in the Metaverse due to the stringent requirements on computing workload, communication throughput, and round-trip latency. In this…

Robotics · Computer Science 2024-07-24 Kan Chen , Zhen Meng , Xiangmin Xu , Changyang She , Philip G. Zhao

With large language models (LLMs) on the rise, in-game interactions are shifting from rigid commands to natural conversations. However, the impacts of LLMs on player performance and game experience remain underexplored. This work explores…

Human-Computer Interaction · Computer Science 2025-07-31 Xin Sun , Lei Wang , Yue Li , Jie Li , Massimo Poesio , Julian Frommel , Koen Hinriks , Jiahuan Pei

World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are often either learned directly from pixels with limited…

Artificial Intelligence · Computer Science 2026-05-26 Minghao Fu , Fan Feng , Nicklas Hansen , Biwei Huang

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects in isolation, as…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yicong Hong , Yiqun Mei , Chongjian Ge , Yiran Xu , Yang Zhou , Sai Bi , Yannick Hold-Geoffroy , Mike Roberts , Matthew Fisher , Eli Shechtman , Kalyan Sunkavalli , Feng Liu , Zhengqi Li , Hao Tan

Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Ling Yang , Bohan Zeng , Jiaming Liu , Hong Li , Minghao Xu , Wentao Zhang , Shuicheng Yan

We introduce GameGen-X, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos. This model facilitates high-quality, open-domain generation by simulating an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Haoxuan Che , Xuanhua He , Quande Liu , Cheng Jin , Hao Chen

When performing complex tasks, humans naturally reason at multiple temporal and spatial resolutions simultaneously. We contend that for an artificially intelligent agent to effectively model human teammates, i.e., demonstrate computational…

Machine Learning · Computer Science 2022-11-15 Liang Zhang , Justin Lieffers , Adarsh Pyarelal

Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce WBench, a comprehensive multi-turn…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Kaining Ying , Hengrui Hu , Siyu Ren , Jiamu Li , Fengjiao Chen , Ziwen Wang , Xuezhi Cao , Xunliang Cai , Henghui Ding

Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Zhaoyang Yang , Yurun Jin , Lizhe Qi , Cong Huang , Kai Chen

Traffic microsimulators are widely used to evaluate road network performance under various ``what-if" conditions. However, the behavior models controlling the actions of the actors are overly simplistic and fails to capture realistic…

Machine Learning · Computer Science 2026-03-20 Yash Ranjan , Rahul Sengupta , Anand Rangarajan , Sanjay Ranka

We explore building generative neural network models of popular reinforcement learning environments. Our world model can be trained quickly in an unsupervised manner to learn a compressed spatial and temporal representation of the…

Machine Learning · Computer Science 2018-05-10 David Ha , Jürgen Schmidhuber

Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling…

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act…

Constructing AI models that respond to text instructions is challenging, especially for sequential decision-making tasks. This work introduces a methodology, inspired by unCLIP, for instruction-tuning generative models of behavior without…

Artificial Intelligence · Computer Science 2024-02-06 Shalev Lifshitz , Keiran Paster , Harris Chan , Jimmy Ba , Sheila McIlraith

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

Robotics · Computer Science 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen