中文
相关论文

相关论文: DINO-WM: World Models on Pre-trained Visual Featur…

200 篇论文

The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could greatly simplify the use of images in any…

Learning robust and generalizable world models is crucial for enabling efficient and scalable robotic control in real-world environments. In this work, we introduce a novel framework for learning world models that accurately capture…

机器人学 · 计算机科学 2025-12-16 Chenhao Li , Andreas Krause , Marco Hutter

Simulation-to-simulation and simulation-to-real world transfer of neural network models have been a difficult problem. To close the reality gap, prior methods to simulation-to-real world transfer focused on domain adaptation, decoupling…

机器学习 · 计算机科学 2020-01-06 Sahika Genc , Sunil Mallya , Sravan Bodapati , Tao Sun , Yunzhe Tao

World models (WMs) demonstrate strong capabilities in prediction, generation, and planning tasks. Existing WMs primarily focus on unstructured data and cannot leverage the ubiquitous structured data, often represented as graphs, in the…

机器学习 · 计算机科学 2025-07-15 Tao Feng , Yexin Wu , Guanyu Lin , Jiaxuan You

Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that…

机器人学 · 计算机科学 2025-09-18 Guanxing Lu , Baoxiong Jia , Puhao Li , Yixin Chen , Ziwei Wang , Yansong Tang , Siyuan Huang

Identifying predictive world models for robots in novel environments from sparse online observations is essential for robot task planning and execution in novel environments. However, existing methods that leverage differentiable…

机器人学 · 计算机科学 2025-05-13 Yifan Zhu , Tianyi Xiang , Aaron Dollar , Zherong Pan

Vision Foundation Models (VFMs) have advanced representation learning through self-supervised methods. However, existing training pipelines are often inflexible, domain-specific, or computationally expensive, which limits their usability…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Mahmut Selman Gokmen , Cody Bumgardner

Model-based planning in robotic domains is challenged by the hybrid nature of physical dynamics, where continuous motion is punctuated by discrete events such as contacts and impacts. Conventional latent world models typically employ…

人工智能 · 计算机科学 2026-05-14 Mingwei Li , Xiaoyuan Zhang , Chengwei Yang , Zilong Zheng , Yaodong Yang

World models for environments with many objects face a combinatorial explosion of states: as the number of objects increases, the number of possible arrangements grows exponentially. In this paper, we learn to generalize over robotic…

机器人学 · 计算机科学 2022-02-14 Ondrej Biza , Thomas Kipf , David Klee , Robert Platt , Jan-Willem van de Meent , Lawson L. S. Wong

World models enable robots to conduct counterfactual reasoning in physical environments by predicting future world states. While conventional approaches often prioritize pixel-level reconstruction of future scenes, such detailed rendering…

机器人学 · 计算机科学 2025-12-22 Zhiwei Zhang , Hui Zhang , Kaihong Huang , Chenghao Shi , Huimin Lu

We introduce Diffusion World Model (DWM), a conditional diffusion model capable of predicting multistep future states and rewards concurrently. As opposed to traditional one-step dynamics models, DWM offers long-horizon predictions in a…

机器学习 · 计算机科学 2024-10-17 Zihan Ding , Amy Zhang , Yuandong Tian , Qinqing Zheng

State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained…

World Action Models (WAMs) have emerged as a promising alternative to Vision-Language-Action (VLA) models for embodied control because they explicitly model how visual observations may evolve under action. Most existing WAMs follow an…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Tianyuan Yuan , Zibin Dong , Yicheng Liu , Hang Zhao

We review current and emerging knowledge-informed and brain-inspired cognitive systems for realizing adversarial defenses, eXplainable Artificial Intelligence (XAI), and zero-shot or few-short learning. Data-driven deep learning models have…

机器学习 · 计算机科学 2024-03-13 Fuseinin Mumuni , Alhassan Mumuni

Operating effectively in novel real-world environments requires robotic systems to estimate and interact with previously unseen objects. Current state-of-the-art models address this challenge by using large amounts of training data and…

机器人学 · 计算机科学 2026-02-06 Octavio Arriaga , Proneet Sharma , Jichen Guo , Marc Otto , Siddhant Kadwe , Rebecca Adam

In this work, we introduce the Time-Aware World Model (TAWM), a model-based approach that explicitly incorporates temporal dynamics. By conditioning on the time-step size, {\Delta}t, and training over a diverse range of {\Delta}t values --…

机器学习 · 计算机科学 2025-06-11 Anh N. Nhu , Sanghyun Son , Ming Lin

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Anurag Bagchi , Zhipeng Bao , Homanga Bharadhwaj , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

World models learn to predict the temporal evolution of visual observations given a control signal, potentially enabling agents to reason about environments through forward simulation. Because of the focus on forward simulation, current…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Yiqing Shen , Aiza Maksutova , Chenjia Li , Mathias Unberath

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations.…

机器人学 · 计算机科学 2025-05-26 Chuning Zhu , Raymond Yu , Siyuan Feng , Benjamin Burchfiel , Paarth Shah , Abhishek Gupta

Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art systems typically rely on modular designs that decouple navigation planning from visual world modeling,…