English

AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving

Computer Vision and Pattern Recognition 2025-10-13 v2 Artificial Intelligence

Abstract

With the rapid advancement of autonomous driving, deploying Vision-Language Models (VLMs) to enhance perception and decision-making has become increasingly common. However, the real-time application of VLMs is hindered by high latency and computational overhead, limiting their effectiveness in time-critical driving scenarios. This challenge is particularly evident when VLMs exhibit over-inference, continuing to process unnecessary layers even after confident predictions have been reached. To address this inefficiency, we propose AD-EE, an Early Exit framework that incorporates domain characteristics of autonomous driving and leverages causal inference to identify optimal exit layers. We evaluate our method on large-scale real-world autonomous driving datasets, including Waymo and the corner-case-focused CODA, as well as on a real vehicle running the Autoware Universe platform. Extensive experiments across multiple VLMs show that our method significantly reduces latency, with maximum improvements reaching up to 57.58%, and enhances object detection accuracy, with maximum gains of up to 44%.

Keywords

Cite

@article{arxiv.2506.05404,
  title  = {AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving},
  author = {Lianming Huang and Haibo Hu and Yufei Cui and Jiacheng Zuo and Shangyu Wu and Nan Guan and Chun Jason Xue},
  journal= {arXiv preprint arXiv:2506.05404},
  year   = {2025}
}

Comments

We believe that the contribution of this paper is not enough, so we integrated it into another new paper. The arXiv ID of the new paper is arXiv:2510.01795