ImagiDrive:自主驾驶的统一想象与规划框架
摘要
自主驾驶需要丰富的上下文理解和精确的预测推理,以安全地导航动态和复杂的环境。视觉语言模型(VLMs)和驾驶世界模型(DWMs)已分别作为强大的解决方案,分别解决了这一挑战的不同方面。VLMs通过理解多模态上下文提供可解释性和强大的行动预测,而DWMs在生成对主动规划至关重要的详细且合理的未来驾驶情景方面表现出色。将VLMs与DWMs集成是一个直观且有前景的策略,却尚未受到充分研究,以充分发挥准确行为预测和真实场景生成的互补优势。然而,这一集成 presents notable challenges, particularly in effectively connecting action-level decisions with high-fidelity pixel-level predictions and maintaining computational efficiency。 In this paper, we propose ImagiDrive, a novel end-to-end autonomous driving framework that integrates a VLM-based driving agent with a DWM-based scene imaginer to form a unified imagination-and-planning loop. The driving agent predicts initial driving trajectories based on multi-modal inputs, guiding the scene imaginer to generate corresponding future scenarios. These imagined scenarios are subsequently utilized to iteratively refine the driving agent's planning decisions. To address efficiency and predictive accuracy challenges inherent in this integration, we introduce an early stopping mechanism and a trajectory selection strategy. Extensive experimental validation on the nuScenes and NAVSIM datasets demonstrates the robustness and superiority of ImagiDrive over previous alternatives under both open-loop and closed-loop conditions。
引用
@article{arxiv.2508.11427,
title = {Bicentric configurations of pentagonal linkages},
author = {Ana Diakvnishvili},
journal= {arXiv preprint arXiv:2508.11427},
year = {2025}
}