English
Related papers

Related papers: RoadscapesQA: A Multitask, Multimodal Dataset for …

200 papers

Understanding novel situations in the traffic domain requires an intricate combination of domain-specific and causal commonsense knowledge. Prior work has provided sufficient perception-based modalities for traffic monitoring, in this…

Computation and Language · Computer Science 2022-12-16 Jiarui Zhang , Filip Ilievski , Aravinda Kollaa , Jonathan Francis , Kaixin Ma , Alessandro Oltramari

Utilizing infrastructure and vehicle-side information to track and forecast the behaviors of surrounding traffic participants can significantly improve decision-making and safety in autonomous driving. However, the lack of real-world…

Computer Vision and Pattern Recognition · Computer Science 2023-05-11 Haibao Yu , Wenxian Yang , Hongzhi Ruan , Zhenwei Yang , Yingjuan Tang , Xu Gao , Xin Hao , Yifeng Shi , Yifeng Pan , Ning Sun , Juan Song , Jirui Yuan , Ping Luo , Zaiqing Nie

Accurate accident anticipation remains challenging when driver cognition and dynamic road conditions are underrepresented in predictive models. In this paper, we propose CAMERA (Context-Aware Multi-modal Enhanced Risk Anticipation), a…

Computational Engineering, Finance, and Science · Computer Science 2025-07-17 Jiaxun Zhang , Haicheng Liao , Yumu Xie , Chengyue Wang , Yanchen Guan , Bin Rao , Zhenning Li

Human perception of the world is shaped by a multitude of viewpoints and modalities. While many existing datasets focus on scene understanding from a certain perspective (e.g. egocentric or third-person views), our dataset offers a panoptic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Hao Chen , Yuqi Hou , Chenyuan Qu , Irene Testini , Xiaohan Hong , Jianbo Jiao

Autonomous vehicles must balance a complex set of objectives. There is no consensus on how they should do so, nor on a model for specifying a desired driving behavior. We created a dataset to help address some of these questions in a…

The complex compositional structure of language makes problems at the intersection of vision and language challenging. But language also provides a strong prior that can result in good superficial performance, without the underlying models…

Computation and Language · Computer Science 2016-04-20 Peng Zhang , Yash Goyal , Douglas Summers-Stay , Dhruv Batra , Devi Parikh

Autonomous vehicles (AV) are expected to reshape future transportation systems, and decision-making is one of the critical modules toward high-level automated driving. To overcome those complicated scenarios that rule-based methods could…

Robotics · Computer Science 2023-09-25 Yuning Wang , Zeyu Han , Yining Xing , Shaobing Xu , Jianqiang Wang

The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Lianqing Zheng , Long Yang , Qunshu Lin , Wenjin Ai , Minghao Liu , Shouyi Lu , Jianan Liu , Hongze Ren , Jingyue Mo , Xiaokai Bai , Jie Bai , Zhixiong Ma , Xichan Zhu

Intelligent Traffic Monitoring (ITMo) technologies hold the potential for improving road safety/security and for enabling smart city infrastructure. Understanding traffic situations requires a complex fusion of perceptual information with…

Computation and Language · Computer Science 2023-07-18 Jiarui Zhang , Filip Ilievski , Kaixin Ma , Aravinda Kollaa , Jonathan Francis , Alessandro Oltramari

The rise of Visual-Language Models (LVLMs) has unlocked new possibilities for seamlessly integrating visual and textual information. However, their ability to interpret cartographic maps remains largely unexplored. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Huy Quang Ung , Guillaume Habault , Yasutaka Nishimura , Hao Niu , Roberto Legaspi , Tomoki Oya , Ryoichi Kojima , Masato Taya , Chihiro Ono , Atsunori Minamikawa , Yan Liu

Amodal perception terms the ability of humans to imagine the entire shapes of occluded objects. This gives humans an advantage to keep track of everything that is going on, especially in crowded situations. Typical perception functions,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-02 Jasmin Breitenstein , Tim Fingscheidt

This paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual Question Answering system up to date, namely the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Ali Furkan Biten , Rubèn Tito , Andres Mafla , Lluis Gomez , Marçal Rusiñol , Minesh Mathew , C. V. Jawahar , Ernest Valveny , Dimosthenis Karatzas

Automatic underground parking has attracted considerable attention as the scope of autonomous driving expands. The auto-vehicle is supposed to obtain the environmental information, track its location, and build a reliable map of the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Jiawei Hou , Qi Chen , Yurong Cheng , Guang Chen , Xiangyang Xue , Taiping Zeng , Jian Pu

Detecting traversable road areas ahead a moving vehicle is a key process for modern autonomous driving systems. A common approach to road detection consists of exploiting color features to classify pixels as road or background. These…

Computer Vision and Pattern Recognition · Computer Science 2014-12-19 Jose M. Alvarez , Theo Gevers , Antonio M. Lopez

Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale vision-language benchmark…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Xinyu Liu , Darryl C. Jacob , Yuxin Liu , Xinsong Du , Muchao Ye , Bolei Zhou , Pan He

Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward…

Traffic accident prediction in driving videos aims to provide an early warning of the accident occurrence, and supports the decision making of safe driving systems. Previous works usually concentrate on the spatial-temporal correlation of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Jianwu Fang , Lei-Lei Li , Kuan Yang , Zhedong Zheng , Jianru Xue , Tat-Seng Chua

Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Zile Guo , Zhan Chen , Enze Zhu , Kan Wei , Yongkang Zou , Xiaoxuan Liu , Lei Wang

We used a 3D simulator to create artificial video data with standardized annotations, aiming to aid in the development of Embodied AI. Our question answering (QA) dataset measures the extent to which a robot can understand human behavior…

Artificial Intelligence · Computer Science 2024-09-18 Takanori Ugai , Kensho Hara , Shusaku Egami , Ken Fukuda

Current datasets for vehicular applications are mostly collected in North America or Europe. Models trained or evaluated on these datasets might suffer from geographical bias when deployed in other regions. Specifically, for scene…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Pedro Azevedo , Emanuella Araújo , Gabriel Pierre , Willams de Lima Costa , João Marcelo Teixeira , Valter Ferreira , Roberto Jones , Veronica Teichrieb