English
Related papers

Related papers: DreamDojo: A Generalist Robot World Model from Lar…

200 papers

Learning a generalist control policy for dexterous manipulation typically relies on large-scale datasets. Given the high cost of real-world data collection, a practical alternative is to generate synthetic data through simulation. However,…

Robotics · Computer Science 2026-03-25 Ruixing Jin , Zicheng Zhu , Ruixiang Ouyang , Sheng Xu , Bo Yue , Zhizheng Wu , Guiliang Liu

We introduce NitroGen, a vision-action foundation model for generalist gaming agents that is trained on 40,000 hours of gameplay videos across more than 1,000 games. We incorporate three key ingredients: 1) an internet-scale video-action…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Loïc Magne , Anas Awadalla , Guanzhi Wang , Yinzhen Xu , Joshua Belofsky , Fengyuan Hu , Joohwan Kim , Ludwig Schmidt , Georgia Gkioxari , Jan Kautz , Yisong Yue , Yejin Choi , Yuke Zhu , Linxi "Jim" Fan

Robot learning methods have the potential for widespread generalization across tasks, environments, and objects. However, these methods require large diverse datasets that are expensive to collect in real-world robotics settings. For robot…

Robotics · Computer Science 2023-02-24 Zoey Chen , Sho Kiami , Abhishek Gupta , Vikash Kumar

Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents and their environments conditioned on multi-modal user…

The field of robotics has made significant strides toward developing generalist robot manipulation policies. However, evaluating these policies in real-world scenarios remains time-consuming and challenging, particularly as the number of…

Robotics · Computer Science 2025-05-27 Yaxuan Li , Yichen Zhu , Junjie Wen , Chaomin Shen , Yi Xu

World models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and…

Artificial Intelligence · Computer Science 2024-10-01 Zhiqi Ge , Hongzhe Huang , Mingze Zhou , Juncheng Li , Guoming Wang , Siliang Tang , Yueting Zhuang

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Anurag Bagchi , Zhipeng Bao , Homanga Bharadhwaj , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constrained to narrow…

Robotics · Computer Science 2026-04-10 Chao Tang , Jiacheng Xu , Haofei Lu , Bolin Zou , Wenlong Dong , Hong Zhang , Danica Kragic

We present LangToMo, a vision-language-action framework structured as a dual-system architecture that uses pixel motion forecasts as intermediate representations. Our high-level System 2, an image diffusion model, generates text-conditioned…

Robotics · Computer Science 2025-08-29 Kanchana Ranasinghe , Xiang Li , E-Ro Nguyen , Cristina Mata , Jongwoo Park , Michael S Ryoo

World models learn general knowledge from videos and simulate experience for training behaviors in imagination, offering a path towards intelligent agents. However, previous world models have been unable to accurately predict object…

Artificial Intelligence · Computer Science 2025-09-30 Danijar Hafner , Wilson Yan , Timothy Lillicrap

Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pre-training is essential for extracting a universal representation. However, current vision-centric pre-training typically relies on either 2D or…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Chen Min , Dawei Zhao , Liang Xiao , Jian Zhao , Xinli Xu , Zheng Zhu , Lei Jin , Jianshu Li , Yulan Guo , Junliang Xing , Liping Jing , Yiming Nie , Bin Dai

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the…

Robotics · Computer Science 2025-07-30 Fangqi Zhu , Hongtao Wu , Song Guo , Yuxiao Liu , Chilam Cheang , Tao Kong

Data scaling has revolutionized fields like natural language processing and computer vision, providing models with remarkable generalization capabilities. In this paper, we investigate whether similar data scaling laws exist in robotics,…

Robotics · Computer Science 2025-10-14 Yingdong Hu , Fanqi Lin , Pingyue Sheng , Chuan Wen , Jiacheng You , Yang Gao

World models are a powerful paradigm in AI and robotics, enabling agents to reason about the future by predicting visual observations or compact latent states. The 1X World Model Challenge introduces an open-source benchmark of real-world…

Machine Learning · Computer Science 2025-10-09 Riccardo Mereu , Aidan Scannell , Yuxin Hou , Yi Zhao , Aditya Jitta , Antonio Dominguez , Luigi Acerbi , Amos Storkey , Paul Chang

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yang Zhou , Hao Shao , Letian Wang , Zhuofan Zong , Hongsheng Li , Steven L. Waslander

Deep reinforcement learning (RL) can acquire complex behaviors from low-level inputs, such as images. However, real-world applications of such methods require generalizing to the vast variability of the real world. Deep networks are known…

Machine Learning · Computer Science 2017-03-13 Chelsea Finn , Tianhe Yu , Justin Fu , Pieter Abbeel , Sergey Levine

Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can come from human videos, which cover diverse real-world…

Humanoid robots, with their human-like form, are uniquely suited for interacting in environments built for people. However, enabling humanoids to reason, plan, and act in complex open-world settings remains a challenge. World models, models…

Robotics · Computer Science 2025-07-10 Muhammad Qasim Ali , Aditya Sridhar , Shahbuland Matiana , Alex Wong , Mohammad Al-Sharman

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics datasets, offering…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haotian Xue , Yipu Chen , Liqian Ma , Zelin Zhao , Lama Moukheiber , Yuchen Zhu , Yongxin Chen

Teaching a multi-fingered dexterous robot to grasp objects in the real world has been a challenging problem due to its high dimensional state and action space. We propose a robot-learning system that can take a small number of human…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Zoey Qiuyu Chen , Karl Van Wyk , Yu-Wei Chao , Wei Yang , Arsalan Mousavian , Abhishek Gupta , Dieter Fox