English
Related papers

Related papers: Towards Unified World Models for Visual Navigation…

200 papers

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics datasets, offering…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haotian Xue , Yipu Chen , Liqian Ma , Zelin Zhao , Lama Moukheiber , Yuchen Zhu , Yongxin Chen

Mobile traffic prediction is a fundamental yet challenging problem for wireless network planning and optimization. Existing models focus on learning static long-term temporal patterns in mobile traffic series, which limits their ability to…

Networking and Internet Architecture · Computer Science 2026-04-10 Xiaoqian Qi , Haoye Chai , Yue Wang , Yong Li

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zuyao Lin , Jianhui Zhang , Peidong Jia , Xiaoguang Zhao , Shanghang Zhang , Xingyu Chen

World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Zeqi Xiao , Yushi Lan , Yifan Zhou , Wenqi Ouyang , Shuai Yang , Yanhong Zeng , Xingang Pan

Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for visual causal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Hao Lu , Ziyang Liu , Guangfeng Jiang , Yuanfei Luo , Sheng Chen , Yangang Zhang , Ying-Cong Chen

This paper introduces UniGen, a novel approach to generating new traffic scenarios for evaluating and improving autonomous driving software through simulation. Our approach models all driving scenario elements in a unified model: the…

Social navigation requires robots to act safely in dynamic human environments. Effective behavior demands thinking ahead: reasoning about how the scene and pedestrians evolve under different robot actions rather than reacting to current…

Robotics · Computer Science 2026-03-19 Tianshuai Hu , Zeying Gong , Lingdong Kong , XiaoDong Mei , Yiyi Ding , Qi Zeng , Ao Liang , Rong Li , Yangyi Zhong , Junwei Liang

Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pre-training is essential for extracting a universal representation. However, current vision-centric pre-training typically relies on either 2D or…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Chen Min , Dawei Zhao , Liang Xiao , Jian Zhao , Xinli Xu , Zheng Zhu , Lei Jin , Jianshu Li , Yulan Guo , Junliang Xing , Liping Jing , Yiming Nie , Bin Dai

In language-guided visual navigation, agents locate target objects in unseen environments using natural language instructions. For reliable navigation in unfamiliar scenes, agents should possess strong perception, planning, and prediction…

Robotics · Computer Science 2025-08-11 Yufeng Zhong , Chengjian Feng , Feng Yan , Fanfan Liu , Liming Zheng , Lin Ma

Spatial reasoning in partially observable environments has often been approached through passive predictive models, yet theories of embodied cognition suggest that genuinely useful representations arise only when perception is tightly…

Artificial Intelligence · Computer Science 2025-04-29 Li Jin , Liu Jia

We study language-conditioned visual navigation (LCVN), in which an embodied agent is asked to follow a natural language instruction based only on an initial egocentric observation. Without access to goal images, the agent must rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yifei Dong , Fengyi Wu , Yilong Dai , Lingdong Kong , Guangyu Chen , Xu Zhu , Qiyu Hu , Tianyu Wang , Johnalbert Garnica , Feng Liu , Siyu Huang , Qi Dai , Zhi-Qi Cheng

Agents operating in complex software environments benefit from reasoning about the consequences of their actions, as even a single incorrect user interface (UI) operation can derail long, artifact-preserving workflows. This challenge is…

We introduce multi-task Visuo-Tactile World Models (VT-WM), which capture the physics of contact through touch reasoning. By complementing vision with tactile sensing, VT-WM better understands robot-object interactions in contact-rich…

Embodied systems experience the world as 'a symphony of flows': a combination of many continuous streams of sensory input coupled to self-motion, interwoven with the dynamics of external objects. These streams obey smooth,…

Machine Learning · Computer Science 2026-01-06 Hansen Jin Lillemark , Benhao Huang , Fangneng Zhan , Yilun Du , Thomas Anderson Keller

We introduce a method for real-time navigation and tracking with differentiably rendered world models. Learning models for control has led to impressive results in robotics and computer games, but this success has yet to be extended to…

Machine Learning · Computer Science 2022-01-26 Baris Kayalibay , Atanas Mirchev , Patrick van der Smagt , Justin Bayer

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of…

A navigable agent needs to understand both high-level semantic instructions and precise spatial perceptions. Building navigation agents centered on Multimodal Large Language Models (MLLMs) demonstrates a promising solution due to their…

Robotics · Computer Science 2026-02-18 Zerui Li , Hongpei Zheng , Fangguo Zhao , Aidan Chan , Jian Zhou , Sihao Lin , Shijie Li , Qi Wu

World-Action Models (WAM) initialized from pre-trained video generation backbones have demonstrated remarkable potential for robot policy learning. However, existing approaches face two critical bottlenecks that hinder performance and…

Embodied agents are expected to perform object navigation in dynamic, open-world environments. However, existing approaches typically rely on static trajectories and a fixed set of object categories during training, overlooking the…

Robotics · Computer Science 2026-04-07 Ming-Ming Yu , Fei Zhu , Wenzhuo Liu , Yirong Yang , Qunbo Wang , Wenjun Wu , Jing Liu

Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured outputs and heterogeneous decoding behaviors, such as autoregressive language generation,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yiwei Zhang , Xuesong Chen , Jin Gao , Hanshi Wang , Fudong Ge , Weiming Hu , Shaoshuai Shi , Zhipeng Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›