English
Related papers

Related papers: AstraNav-World: World Model for Foresight Control …

200 papers

Embodied agents often struggle with efficient navigation because they rely primarily on partial egocentric observations, which restrict global foresight and lead to inefficient exploration. In contrast, humans plan using maps: we reason…

Robotics · Computer Science 2026-02-19 Yuzhuo Ao , Anbang Wang , Yu-Wing Tai , Chi-Keung Tang

Action-conditioned robot world models generate future video frames of the manipulated scene given a robot action sequence, offering a promising alternative for simulating tasks that are difficult to model with traditional physics engines.…

Robotics · Computer Science 2026-03-27 Jai Bardhan , Patrik Drozdik , Josef Sivic , Vladimir Petrik

Deploying autonomous agents in real world environments is challenging, particularly for navigation, where systems must adapt to situations they have not encountered before. Traditional learning approaches require substantial amounts of…

Robotics · Computer Science 2026-03-10 Quang-Anh N. D. , Duc Pham , Minh-Anh Nguyen , Tung Doan , Tuan Dang

End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning cannot anticipate future scene evolution, often leading to…

Robotics · Computer Science 2026-04-29 Chuyao Fu , Shengzhe Gan , Zhuoli Ouyang , Yuhan Rui , Xiaowei Chi , Sirui Han , Jiankun Wang , Hong Zhang

Robust autonomous navigation for Autonomous Aerial Vehicles (AAVs) in complex environments is a critical capability. However, modern end-to-end navigation faces a key challenge: the high-frequency control loop needed for agile flight…

Robotics · Computer Science 2026-03-10 Yude Li , Zhexuan Zhou , Huizhe Li , Youmin Gong , Jie Mei

Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning from simulation to real-world deployment, primarily due to…

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Robotic imitation learning has advanced from solving static tasks to addressing dynamic interaction scenarios, but testing and evaluation remain costly and challenging due to the need for real-time interaction with dynamic environments. We…

Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing approaches are often slow and struggle to capture physically consistent interactions over…

Video-based world models offer a powerful paradigm for embodied simulation and planning, yet state-of-the-art models often generate physically implausible manipulations - such as object penetration and anti-gravity motion - due to training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yuzhi Chen , Ronghan Chen , Dongjie Huo , Yandan Yang , Dekang Qi , Haoyun Liu , Tong Lin , Shuang Zeng , Junjin Xiao , Xinyuan Chang , Feng Xiong , Xing Wei , Zhiheng Ma , Mu Xu

Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Bozhou Zhang , Nan Song , Yuang Wang , Jiankang Deng , Xiatian Zhu , Li Zhang

World models learn the consequences of actions in vision-based interactive systems. However, in practical scenarios such as autonomous driving, there commonly exists noncontrollable dynamics independent of the action signals, making it…

Machine Learning · Computer Science 2022-12-20 Minting Pan , Xiangming Zhu , Yunbo Wang , Xiaokang Yang

Learning predictive world models from raw visual observations is a central challenge in reinforcement learning (RL), especially for robotics and continuous control. Conventional model-based RL frameworks directly condition future…

Robotics · Computer Science 2026-03-13 Jseen Zhang , Gabriel Adineera , Jinzhou Tan , Jinoh Kim

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jianbiao Mei , Yu Yang , Xuemeng Yang , Licheng Wen , Jiajun Lv , Botian Shi , Yong Liu

General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning, which primarily focuses on optimal actions, a world…

Machine Learning · Computer Science 2026-04-03 Yuejiang Liu , Fan Feng , Lingjing Kong , Weifeng Lu , Jinzhou Tang , Kun Zhang , Kevin Murphy , Chelsea Finn , Yilun Du

Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned…

Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geometry, effectively anticipating the future requires capturing…

Robotics · Computer Science 2026-03-12 Xiaoxu Xu , Hao Li , Jinhui Ye , Yilun Chen , Jia Zeng , Xinyi Chen , Linning Xu , Dahua Lin , Weixin Li , Jiangmiao Pang

Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly transferring these…

Artificial Intelligence · Computer Science 2026-03-03 Qianqian Bai , Zhongpu Chen , Ling Luo , Huaming Du , Yuqian Lei , Ziyun Jiao

Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-world tasks. However, scaling these systems remains difficult because collecting…

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Fuxiang Yang , Donglin Di , Lulu Tang , Xuancheng Zhang , Lei Fan , Hao Li , Chen Wei , Tonghua Su , Baorui Ma