English
Related papers

Related papers: RoadscapesQA: A Multitask, Multimodal Dataset for …

200 papers

For an autonomous vehicle, situation understand-ing is a key capability towards safe and comfortable decision-making and navigation. Information is in general provided bymultiple sources. Prior information about the road topology andtraffic…

Robotics · Computer Science 2021-10-25 Corentin Sanchez , Philippe Xu , Alexandre Armand , Philippe Bonnifait

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Understanding the complex, multi-agent dynamics of urban traffic remains a fundamental challenge for video language models. This paper introduces Urban Dynamics VideoQA, a benchmark dataset that captures the unscripted real-world behavior…

Safety on roads is of uttermost importance, especially in the context of autonomous vehicles. A critical need is to detect and communicate disruptive incidents early and effectively. In this paper we propose a system based on an…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Alex Levering , Martin Tomko , Devis Tuia , Kourosh Khoshelham

Representing diverse and plausible future trajectories is critical for motion forecasting in autonomous driving. However, efficiently capturing these trajectories in a compact set remains challenging. This study introduces a novel approach…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Abhishek Vivekanandan , J. Marius Zöllner

Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes. It…

Computation and Language · Computer Science 2018-09-05 Semih Yagcioglu , Aykut Erdem , Erkut Erdem , Nazli Ikizler-Cinbis

Realistic scene reconstruction and view synthesis are essential for advancing autonomous driving systems by simulating safety-critical scenarios. 3D Gaussian Splatting excels in real-time rendering and static scene reconstructions but…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Mustafa Khan , Hamidreza Fazlali , Dhruv Sharma , Tongtong Cao , Dongfeng Bai , Yuan Ren , Bingbing Liu

Visual question answering is an important task in both natural language and vision understanding. However, in most of the public visual question answering datasets such as VQA, CLEVR, the questions are human generated that specific to the…

Computation and Language · Computer Science 2022-08-08 Bingning Wang , Feiyang Lv , Ting Yao , Yiming Yuan , Jin Ma , Yu Luo , Haijin Liang

Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their application to autonomous driving. However, high-level reasoning in safety-critical scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Seungjun Yu , Seonho Lee , Namho Kim , Jaeyo Shin , Junsung Park , Wonjeong Ryu , Raehyuk Jung , Hyunjung Shim

Accurate speed estimation of road vehicles is important for several reasons. One is speed limit enforcement, which represents a crucial tool in decreasing traffic accidents and fatalities. Compared with other research areas and domains, the…

Machine Learning · Computer Science 2022-12-06 Slobodan Djukanović , Nikola Bulatović , Ivana Čavor

Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsically black box models, which lack transparency and safety in autonomous driving. To make…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Maryam Sadat Hosseini Azad , Shahriar Baradaran Shokouhi

Accurate driving behavior recognition and reasoning are critical for autonomous driving video understanding. However, existing methods often tend to dig out the shallow causal, fail to address spurious correlations across modalities, and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Tongtong Cheng , Rongzhen Li , Yixin Xiong , Tao Zhang , Jing Wang , Kai Liu

Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-world applications could be developed by integrating…

Accurate lane detection is essential for automated driving, enabling safe and reliable vehicle navigation across a variety of road scenarios. Numerous datasets have been introduced to support the development and evaluation of lane detection…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jörg Gamerdinger , Sven Teufel , Oliver Bringmann

This paper offers openly available microscopic vehicle trajectory (MVT) datasets collected using unmanned aerial vehicles (UAVs) in heterogeneous, area-based urban traffic conditions. Traditional roadside video collection often fails in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Yawar Ali , K. Ramachandra Rao , Ashish Bhaskar , Niladri Chatterjee

Recently, road scene-graph representations used in conjunction with graph learning techniques have been shown to outperform state-of-the-art deep learning techniques in tasks including action classification, risk assessment, and collision…

Computer Vision and Pattern Recognition · Computer Science 2022-01-03 Arnav Vaibhav Malawade , Shih-Yuan Yu , Brandon Hsu , Harsimrat Kaeley , Anurag Karra , Mohammad Abdullah Al Faruque

We introduce RoScenes, the largest multi-view roadside perception dataset, which aims to shed light on the development of vision-centric Bird's Eye View (BEV) approaches for more challenging traffic scenes. The highlights of RoScenes…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Xiaosu Zhu , Hualian Sheng , Sijia Cai , Bing Deng , Shaopeng Yang , Qiao Liang , Ken Chen , Lianli Gao , Jingkuan Song , Jieping Ye

Most approaches for instance-aware semantic labeling traditionally focus on accuracy. Other aspects like runtime and memory footprint are arguably as important for real-time applications such as autonomous driving. Motivated by this…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 Davy Neven , Bert De Brabandere , Stamatios Georgoulis , Marc Proesmans , Luc Van Gool

Visual question answering (or VQA) is a new and exciting problem that combines natural language processing and computer vision techniques. We present a survey of the various datasets and models that have been used to tackle this task. The…

Computation and Language · Computer Science 2017-05-12 Akshay Kumar Gupta

Road unevenness significantly impacts the safety and comfort of traffic participants, especially vulnerable groups such as cyclists and wheelchair users. To train models for comprehensive road surface assessments, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Alexandra Kapp , Edith Hoffmann , Esther Weigmann , Helena Mihaljević