English
Related papers

Related papers: Multimodal Framework for Explainable Autonomous Dr…

200 papers

Concept bottleneck models have been successfully used for explainable machine learning by encoding information within the model with a set of human-defined concepts. In the context of human-assisted or autonomous driving, explainability…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Jessica Echterhoff , An Yan , Kyungtae Han , Amr Abdelraouf , Rohit Gupta , Julian McAuley

In today's world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources…

Recent advancements in vision foundation models (VFMs) have revolutionized visual perception in 2D, yet their potential for 3D scene understanding, particularly in autonomous driving applications, remains underexplored. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Lingdong Kong , Xiang Xu , Youquan Liu , Jun Cen , Runnan Chen , Wenwei Zhang , Liang Pan , Kai Chen , Ziwei Liu

While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce OmniAD, a novel framework that unifies anomaly detection and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Shifang Zhao , Yiheng Lin , Lu Han , Yao Zhao , Yunchao Wei

In autonomous driving, perception systems are piv otal as they interpret sensory data to understand the envi ronment, which is essential for decision-making and planning. Ensuring the safety of these perception systems is fundamental for…

Robotics · Computer Science 2024-11-19 Urvishkumar Bharti , Vikram Shahapur

Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning and open-world generalization. However, the excessive computational overhead and high…

Robotics · Computer Science 2026-05-26 Ruoyu Yao , Ruiguo Zhong , Pei Liu , Mingxing Peng , Rui Yang , Jun Ma

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack the ability to support…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Guile Wu , David Huang , Dongfeng Bai , Bingbing Liu

Today's autonomous vehicles rely extensively on high-definition 3D maps to navigate the environment. While this approach works well when these maps are completely up-to-date, safe autonomous vehicles must be able to corroborate the map's…

Computer Vision and Pattern Recognition · Computer Science 2016-12-09 Ari Seff , Jianxiong Xiao

Pedestrian crossing intention prediction is essential for the deployment of autonomous vehicles (AVs) in urban environments. Ideal prediction provides AVs with critical environmental cues, thereby reducing the risk of pedestrian-related…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yuanzhe Li , Steffen Müller

To plan safe maneuvers and act with foresight, autonomous vehicles must be capable of accurately predicting the uncertain future. In the context of autonomous driving, deep neural networks have been successfully applied to learning…

Robotics · Computer Science 2022-08-02 Salar Arbabi , Davide Tavernini , Saber Fallah , Richard Bowden

Autonomous vehicles need to accomplish their tasks while interacting with human drivers in traffic. It is thus crucial to equip autonomous vehicles with artificial reasoning to better comprehend the intentions of the surrounding traffic,…

Artificial Intelligence · Computer Science 2023-11-02 Xiao Li , Kaiwen Liu , H. Eric Tseng , Anouck Girard , Ilya Kolmanovsky

In mixed-traffic environments, autonomous vehicles (AVs) must interact with heterogeneous human-driven vehicles (HVs) whose intentions and driving styles vary across individuals and scenarios. Such variability introduces uncertainty into…

Robotics · Computer Science 2026-03-18 Xiaoyun Qiu , Haichao Liu , Yue Pan , Jun Ma , Xinhu Zheng

Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yi Xu , Yuxin Hu , Zaiwei Zhang , Gregory P. Meyer , Siva Karthik Mustikovela , Siddhartha Srinivasa , Eric M. Wolff , Xin Huang

In the field of autonomous driving, end-to-end deep learning models show great potential by learning driving decisions directly from sensor data. However, training these models requires large amounts of labeled data, which is time-consuming…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Wenhao Jiang , Duo Li , Menghan Hu , Chao Ma , Ke Wang , Zhipeng Zhang

Given the wide adoption of multimodal sensors (e.g., camera, lidar, radar) by autonomous vehicles (AVs), deep analytics to fuse their outputs for a robust perception become imperative. However, existing fusion methods often make two…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Pengfei Hu , Yuhang Qian , Tianyue Zheng , Ang Li , Zhe Chen , Yue Gao , Xiuzhen Cheng , Jun Luo

The comprehensiveness of vehicle-to-everything (V2X) recognition enriches and holistically shapes the global Birds-Eye-View (BEV) perception, incorporating rich semantics and integrating driving scene information, thereby serving features…

Robotics · Computer Science 2024-04-23 Fukang Li , Wenlin Ou , Kunpeng Gao , Yuwen Pang , Yifei Li , Henry Fan

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Tianwen Qian , Jingjing Chen , Linhai Zhuo , Yang Jiao , Yu-Gang Jiang

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Building multimodal dialogue understanding capabilities situated in the in-cabin context is crucial to enhance passenger comfort in autonomous vehicle (AV) interaction systems. To this end, understanding passenger intents from spoken…

Computation and Language · Computer Science 2020-07-09 Eda Okur , Shachi H Kumar , Saurav Sahay , Lama Nachman

Current research in semantic bird's-eye view segmentation for autonomous driving focuses solely on optimizing neural network models using a single dataset, typically nuScenes. This practice leads to the development of highly specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Manuel Alejandro Diaz-Zapata , Wenqian Liu , Robin Baruffa , Christian Laugier