中文
相关论文

相关论文: Uni-World VLA: Interleaved World Modeling and Plan…

200 篇论文

Vision-Language-Action (VLA) models have demonstrated potential in autonomous driving. However, two critical challenges hinder their development: (1) Existing VLA architectures are typically based on imitation learning in open-loop setup…

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action,…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Haoyu Zhen , Xiaowen Qiu , Peihao Chen , Jincheng Yang , Xin Yan , Yilun Du , Yining Hong , Chuang Gan

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Julong Wei , Shanshuai Yuan , Pengfei Li , Qingda Hu , Zhongxue Gan , Wenchao Ding

Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Xiyang Wang , Xinlin Wang , Tingguang Zhou , Gong Chen , Xingtai Gui , Zhi Xu , Xiaolei Wu , Feiyang Tan , Hangning Zhou , Mu Yang

Autonomous driving systems face significant challenges in achieving human-like adaptability, robustness, and interpretability in complex, open-world environments. These challenges stem from fragmented architectures, limited generalization…

机器人学 · 计算机科学 2025-08-01 Yi Zhang , Erik Leo Haß , Kuo-Yi Chao , Nenad Petrovic , Yinglei Song , Chengdong Wu , Alois Knoll

Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck for large-scale deployment. To address this challenge, some works use LLMs and VLMs for…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Hao Shao , Letian Wang , Yang Zhou , Yuxuan Hu , Zhuofan Zong , Steven L. Waslander , Wei Zhan , Hongsheng Li

Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-horizon tasks that require defined goal states, such as LEGO…

Typically, autonomous driving adopts a modular design, which divides the full stack into perception, prediction, planning and control parts. Though interpretable, such modular design tends to introduce a substantial amount of redundancy.…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Fan Jia , Weixin Mao , Yingfei Liu , Yucheng Zhao , Yuqing Wen , Chi Zhang , Xiangyu Zhang , Tiancai Wang

General vision-language models increasingly support unified spatiotemporal reasoning over long video streams, yet deploying such capabilities on robots remains constrained by the prohibitive latency of processing long-horizon histories and…

机器人学 · 计算机科学 2026-02-19 Jingjing Fan , Yushan Liu , Shoujie Li , Botao Ren , Siyuan Li , Xiao-Ping Zhang , Wenbo Ding , Zhidong Deng

We present OpenDriveVLA, a Vision Language Action model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially grounded driving actions by leveraging multimodal inputs,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Xingcheng Zhou , Xuyuan Han , Feng Yang , Yunpu Ma , Volker Tresp , Alois Knoll

Various world model frameworks are being developed today based on autoregressive frameworks that rely on discrete representations of actions and observations, and these frameworks are succeeding in constructing interactive generative models…

机器学习 · 计算机科学 2025-03-14 Kohei Hayashi , Masanori Koyama , Julian Jorge Andrade Guerreiro

We present an integrated approach for perception and control for an autonomous vehicle and demonstrate this approach in a high-fidelity urban driving simulator. Our approach first builds a model for the environment, then trains a policy…

系统与控制 · 电气工程与系统科学 2020-03-19 Ali Baheri , Ilya Kolmanovsky , Anouck Girard , H. Eric Tseng , Dimitar Filev

Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extension of this journey beyond Vision-to-Action (VA). In practice, driving VLAs have often…

Recent vision-language-action (VLA) models have significantly advanced robotic manipulation by unifying perception, reasoning, and control. To achieve such integration, recent studies adopt a predictive paradigm that models future visual…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yijie Zhu , Jie He , Rui Shao , Kaishen Yuan , Tao Tan , Xiaochen Yuan , Zitong Yu

Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite…

Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geometry, effectively anticipating the future requires capturing…

机器人学 · 计算机科学 2026-03-12 Xiaoxu Xu , Hao Li , Jinhui Ye , Yilun Chen , Jia Zeng , Xinyi Chen , Linning Xu , Dahua Lin , Weixin Li , Jiangmiao Pang

Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Dapeng Zhang , Zhenlong Yuan , Zhangquan Chen , Chih-Ting Liao , Yinda Chen , Fei Shen , Qingguo Zhou , Tat-Seng Chua

Scene understanding and risk-aware attentions are crucial for human drivers to make safe and effective driving decisions. To imitate this cognitive ability in urban autonomous driving while ensuring the transparency and interpretability, we…

机器人学 · 计算机科学 2025-07-22 Haichao Liu , Haoren Guo , Pei Liu , Benshan Ma , Yuxiang Zhang , Jun Ma , Tong Heng Lee

We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model and world model in one single framework. The world model…

机器人学 · 计算机科学 2025-06-27 Jun Cen , Chaohui Yu , Hangjie Yuan , Yuming Jiang , Siteng Huang , Jiayan Guo , Xin Li , Yibing Song , Hao Luo , Fan Wang , Deli Zhao , Hao Chen