English

Talk2BEV: Language-enhanced Bird's-eye View Maps for Autonomous Driving

Computer Vision and Pattern Recognition 2023-11-15 v2 Robotics

Abstract

Talk2BEV is a large vision-language model (LVLM) interface for bird's-eye view (BEV) maps in autonomous driving contexts. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set of object categories and driving scenarios, Talk2BEV blends recent advances in general-purpose language and vision models with BEV-structured map representations, eliminating the need for task-specific models. This enables a single system to cater to a variety of autonomous driving tasks encompassing visual and spatial reasoning, predicting the intents of traffic actors, and decision-making based on visual cues. We extensively evaluate Talk2BEV on a large number of scene understanding tasks that rely on both the ability to interpret free-form natural language queries, and in grounding these queries to the visual context embedded into the language-enhanced BEV map. To enable further research in LVLMs for autonomous driving scenarios, we develop and release Talk2BEV-Bench, a benchmark encompassing 1000 human-annotated BEV scenarios, with more than 20,000 questions and ground-truth responses from the NuScenes dataset.

Keywords

Cite

@article{arxiv.2310.02251,
  title  = {Talk2BEV: Language-enhanced Bird's-eye View Maps for Autonomous Driving},
  author = {Tushar Choudhary and Vikrant Dewangan and Shivam Chandhok and Shubham Priyadarshan and Anushka Jain and Arun K. Singh and Siddharth Srivastava and Krishna Murthy Jatavallabhula and K. Madhava Krishna},
  journal= {arXiv preprint arXiv:2310.02251},
  year   = {2023}
}

Comments

Project page at https://llmbev.github.io/talk2bev/

R2 v1 2026-06-28T12:39:41.658Z