中文
相关论文

相关论文: World Guidance: World Modeling in Condition Space …

200 篇论文

Video generation models (VGMs) have received extensive attention recently and serve as promising candidates for general-purpose large vision models. While they can only generate short videos each time, existing methods achieve long video…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Yuanhui Huang , Wenzhao Zheng , Yuan Gao , Xin Tao , Pengfei Wan , Di Zhang , Jie Zhou , Jiwen Lu

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates errors, causing…

机器人学 · 计算机科学 2026-03-16 Minghao Jin , Mozheng Liao , Mingfei Han , Zhihui Li , Xiaojun Chang

Vision-Language-Action (VLA) models convert high-level language instructions into concrete, executable actions, a task that is especially challenging in open-world environments. We present Visual Foresight Planning (ForeAct), a general and…

机器人学 · 计算机科学 2026-02-16 Zhuoyang Zhang , Shang Yang , Qinghao Hu , Luke J. Huang , James Hou , Yufei Sun , Yao Lu , Song Han

Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Sicheng Zuo , Yuxuan Li , Wenzhao Zheng , Zheng Zhu , Jie Zhou , Jiwen Lu

General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning, which primarily focuses on optimal actions, a world…

机器学习 · 计算机科学 2026-04-03 Yuejiang Liu , Fan Feng , Lingjing Kong , Weifeng Lu , Jinzhou Tang , Kun Zhang , Kevin Murphy , Chelsea Finn , Yilun Du

Vision-Language-Action (VLA) models generalize semantically well but often lack fine-grained modeling of world dynamics. We present MotuBrain, a unified World Action Model that jointly models video and action under a UniDiffuser formulation…

This paper presents the World-Action Model (WAM), an action-regularized world model that jointly reasons over future visual observations and the actions that drive state transitions. Unlike conventional world models trained solely via image…

人工智能 · 计算机科学 2026-04-01 Yuci Han , Alper Yilmaz

Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and decoupled from trajectory planning. While recent efforts aim…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Zhida Zhao , Talas Fu , Yifan Wang , Lijun Wang , Huchuan Lu

World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: the "Imagine-then-Execute" approach, which uses video prediction to infer actions…

机器人学 · 计算机科学 2026-05-12 Qiuxuan Feng , Jiale Yu , Jiaming Liu , Yueru Jia , Zhuangzhe Wu , Hao Chen , Zezhong Qian , Shuo Gu , Peng Jia , Siwei Ma , Shanghang Zhang

Evaluating robot control policies is difficult: real-world testing is costly, and handcrafted simulators require manual effort to improve in realism and generality. We propose a world-model-based policy evaluation environment (WorldGym), an…

机器人学 · 计算机科学 2025-10-01 Julian Quevedo , Ansh Kumar Sharma , Yixiang Sun , Varad Suryavanshi , Percy Liang , Sherry Yang

World models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and…

人工智能 · 计算机科学 2024-10-01 Zhiqi Ge , Hongzhe Huang , Mingze Zhou , Juncheng Li , Guoming Wang , Siliang Tang , Yueting Zhuang

Learning new skills by observing humans' behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans' decision-making processes, focusing on learning a model to plan goal-directed actions in…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Jing Bi , Jiebo Luo , Chenliang Xu

State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained…

The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned World Models as generative simulators, enabling policy optimization entirely within…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Jiaxuan Gao , Yongjian Guo , Zhong Guan , Wen Huang , Wanlun Ma , Xi Xiao , Junwu Xiong , Sheng Wen

Virtual and augmented reality systems increasingly demand intelligent adaptation to user behaviors for enhanced interaction experiences. Achieving this requires accurately understanding human intentions and predicting future situated…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Yuan Xu , Zimu Zhang , Xiaoxuan Ma , Wentao Zhu , Yu Qiao , Yizhou Wang

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through photorealistic environments following natural-language instructions. One main obstacle existing in VLN is data scarcity, leading to poor…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Yu Zhong , Rui Zhang , Zihao Zhang , Shuo Wang , Chuan Fang , Xishan Zhang , Jiaming Guo , Shaohui Peng , Di Huang , Yanyang Yan , Xing Hu , Qi Guo

Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Xiaodong Mei , Diankun Zhang , Hongwei Xie , Guang Chen , Hangjun Ye , Dan Xu

End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on expert demonstrations. However, imitation learning inherently…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Zihao Sheng , Xin Ye , Jingru Luo , Sikai Chen , Liu Ren

World Action Models (WAMs) have emerged as a promising alternative to Vision-Language-Action (VLA) models for embodied control because they explicitly model how visual observations may evolve under action. Most existing WAMs follow an…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Tianyuan Yuan , Zibin Dong , Yicheng Liu , Hang Zhao