English

LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset

Computer Vision and Pattern Recognition 2026-04-07 v2 Robotics

Abstract

In real-world domains such as self-driving, generalization to rare scenarios remains a fundamental challenge. To address this, we introduce a new dataset designed for end-to-end driving that focuses on long-tail driving events. We provide multi-view video data, trajectories, high-level instructions, and detailed reasoning traces, facilitating in-context learning and few-shot generalization. The resulting benchmark for multimodal models, such as VLMs and VLAs, goes beyond safety and comfort metrics by evaluating instruction following and semantic coherence between model outputs. The multilingual reasoning traces in English, Spanish, and Chinese are from domain experts with diverse cultural backgrounds. Thus, our dataset is a unique resource for studying how different forms of reasoning affect driving competence. Our dataset is available at: https://hf.co/datasets/kit-mrt/kitscenes-longtail

Keywords

Cite

@article{arxiv.2603.23607,
  title  = {LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset},
  author = {Royden Wagner and Omer Sahin Tas and Jaime Villa and Felix Hauser and Yinzhe Shen and Marlon Steiner and Dominik Strutz and Carlos Fernandez and Christian Kinzig and Guillermo S. Guitierrez-Cabello and Hendrik Königshof and Fabian Immel and Richard Schwarzkopf and Nils Alexander Rack and Kevin Rösch and Kaiwen Wang and Jan-Hendrik Pauls and Martin Lauer and Igor Gilitschenski and Holger Caesar and Christoph Stiller},
  journal= {arXiv preprint arXiv:2603.23607},
  year   = {2026}
}

Comments

21 pages; v2: update MMS values (bugfix)