中文
相关论文

相关论文: RoadscapesQA: A Multitask, Multimodal Dataset for …

200 篇论文

The future of automated driving (AD) is rooted in the development of robust, fair and explainable artificial intelligence methods. Upon request, automated vehicles must be able to explain their decisions to the driver and the car…

人工智能 · 计算机科学 2023-08-25 Nassim Belmecheri , Arnaud Gotlieb , Nadjib Lazaar , Helge Spieker

Advances in perception for self-driving cars have accelerated in recent years due to the availability of large-scale datasets, typically collected at specific locations and under nice weather conditions. Yet, to achieve the high safety…

Recent advancements in generative models have provided promising solutions for synthesizing realistic driving videos, which are crucial for training autonomous driving perception models. However, existing approaches often struggle with…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Wei Wu , Xi Guo , Weixuan Tang , Tingxuan Huang , Chiyu Wang , Dongyue Chen , Chenjing Ding

Maintaining situational awareness in complex driving scenarios is challenging. It requires continuously prioritizing attention among extensive scene entities and understanding how prominent hazards might affect the ego vehicle. While…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Yaoqi Huang , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Categorizing driving scenes via visual perception is a key technology for safe driving and the downstream tasks of autonomous vehicles. Traditional methods infer scene category by detecting scene-related objects or using a classifier that…

机器人学 · 计算机科学 2021-03-11 Shaochi Hu , Hanwei Fan , Biao Gao , XijunZhao , Huijing Zhao

Real-world aerial scene understanding is limited by a lack of datasets that contain densely annotated images curated under a diverse set of conditions. Due to inherent challenges in obtaining such images in controlled real-world settings,…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Sahil Khose , Anisha Pal , Aayushi Agarwal , Deepanshi , Judy Hoffman , Prithvijit Chattopadhyay

Safe highway autonomy for heavy trucks remains an open and unsolved challenge: due to long braking distances, scene understanding of hundreds of meters is required for anticipatory planning and to allow safe braking margins. However,…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Filippo Ghilotti , Edoardo Palladin , Samuel Brucker , Adam Sigal , Mario Bijelic , Felix Heide

Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets have been proposed for developing document VQA systems,…

计算与语言 · 计算机科学 2023-01-13 Ryota Tanaka , Kyosuke Nishida , Kosuke Nishida , Taku Hasegawa , Itsumi Saito , Kuniko Saito

Aerial scene recognition is a fundamental research problem in interpreting high-resolution aerial imagery. Over the past few years, most studies focus on classifying an image into one scene category, while in real-world scenarios, it is…

计算机视觉与模式识别 · 计算机科学 2022-02-16 Yuansheng Hua , Lichao Mou , Pu Jin , Xiao Xiang Zhu

Recent advances in multi-modal large language models (MLLMs) have demonstrated strong performance across various domains; however, their ability to comprehend driving scenes remains less proven. The complexity of driving scenarios, which…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Sung-Yeon Park , Can Cui , Yunsheng Ma , Ahmadreza Moradipari , Rohit Gupta , Kyungtae Han , Ziran Wang

Predicting the interaction between pedestrian and vehicle is essential for autonomous driving safety in unstructured and semi-structured scenarios; however, this task is severely hindered by the scarcity of public datasets that feature…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Haoyang Peng , Qian Hu , Songan Zhang , Ming Yang

Autonomous driving is becoming a reality, yet vehicles still need to rely on complex sensor fusion to understand the scene they act in. The ability to discern static environment and dynamic entities provides a comprehension of the road…

计算机视觉与模式识别 · 计算机科学 2018-11-21 Lorenzo Berlincioni , Federico Becattini , Leonardo Galteri , Lorenzo Seidenari , Alberto Del Bimbo

We present the Qualitative Explainable Graph (QXG): a unified symbolic and qualitative representation for scene understanding in urban mobility. QXG enables the interpretation of an automated vehicle's environment using sensor data and…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Nassim Belmecheri , Arnaud Gotlieb , Nadjib Lazaar , Helge Spieker

Intelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle's driving behavior for safe and efficient navigation. This is particularly critical in developing…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Chirag Parikh , Rohit Saluja , C. V. Jawahar , Ravi Kiran Sarvadevabhatla

Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected…

计算机视觉与模式识别 · 计算机科学 2021-10-26 Jiageng Mao , Minzhe Niu , Chenhan Jiang , Hanxue Liang , Jingheng Chen , Xiaodan Liang , Yamin Li , Chaoqiang Ye , Wei Zhang , Zhenguo Li , Jie Yu , Hang Xu , Chunjing Xu

It's important to monitor road issues such as bumps and potholes to enhance safety and improve road conditions. Smartphones are equipped with various built-in sensors that offer a cost-effective and straightforward way to assess road…

Foundation models have indeed made a profound impact on various fields, emerging as pivotal components that significantly shape the capabilities of intelligent systems. In the context of intelligent vehicles, leveraging the power of…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Sheng Luo , Wei Chen , Wanxin Tian , Rui Liu , Luanxuan Hou , Xiubao Zhang , Haifeng Shen , Ruiqi Wu , Shuyi Geng , Yi Zhou , Ling Shao , Yi Yang , Bojun Gao , Qun Li , Guobin Wu

An accurate understanding of a self-driving vehicle's surrounding environment is crucial for its navigation system. To enhance the effectiveness of existing algorithms and facilitate further research, it is essential to provide…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Abtin Mahyar , Hossein Motamednia , Dara Rahmati

3D visual grounding aims to localize the object in 3D point cloud scenes that semantically corresponds to given natural language sentences. It is very critical for roadside infrastructure system to interpret natural languages and localize…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Panquan Yang , Junfei Huang , Zongzhangbao Yin , Yingsong Hu , Anni Xu , Xinyi Luo , Xueqi Sun , Hai Wu , Sheng Ao , Zhaoxing Zhu , Chenglu Wen , Cheng Wang

Image captioning is a computer vision task that involves generating natural language descriptions for images. This method has numerous applications in various domains, including image retrieval systems, medicine, and various industries.…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Sai Suprabhanu Nallapaneni , Subrahmanyam Konakanchi