中文
相关论文

相关论文: DriveXQA: Cross-modal Visual Question Answering fo…

200 篇论文

We present a Collaborative Agent-Based Framework for Multi-Image Reasoning. Our approach tackles the challenge of interleaved multimodal reasoning across diverse datasets and task formats by employing a dual-agent system: a language-based…

Existing benchmarks for Vision-Language Model (VLM) on autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities…

计算与语言 · 计算机科学 2025-03-28 Yue Li , Meng Tian , Zhenyu Lin , Jiangtong Zhu , Dechang Zhu , Haiqiang Liu , Zining Wang , Yueyi Zhang , Zhiwei Xiong , Xinhai Zhao

Multi-modal large language models (MLLMs) have demonstrated remarkable vision-language capabilities, primarily due to the exceptional in-context understanding and multi-task learning strengths of large language models (LLMs). The advent of…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Jianing Li , Xi Nan , Ming Lu , Li Du , Shanghang Zhang

Autonomous driving systems face significant challenges in achieving human-like adaptability, robustness, and interpretability in complex, open-world environments. These challenges stem from fragmented architectures, limited generalization…

机器人学 · 计算机科学 2025-08-01 Yi Zhang , Erik Leo Haß , Kuo-Yi Chao , Nenad Petrovic , Yinglei Song , Chengdong Wu , Alois Knoll

Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this work, we focus on investigating the advantages and…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Xuejing Liu , Wei Tang , Xinzhe Ni , Jinghui Lu , Rui Zhao , Zechao Li , Fei Tan

Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the existence of specific visual abilities, it is easier for the…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Manu Gaur , Darshan Singh S , Makarand Tapaswi

The rapid development of Vision-Language models (VLMs) and Multimodal Language Models (MLLMs) in autonomous driving research has significantly reshaped the landscape by enabling richer scene understanding, context-aware reasoning, and more…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Karthik Mohan , Sonam Singh , Amit Arvind Kale

We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o,…

机器人学 · 计算机科学 2024-10-29 Shiva Sreeram , Tsun-Hsuan Wang , Alaa Maalouf , Guy Rosman , Sertac Karaman , Daniela Rus

Vision-Language-Action (VLA) models have recently emerged in autonomous driving, with the promise of leveraging rich world knowledge to improve the cognitive capabilities of driving systems. However, adapting such models for driving tasks…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Yongkang Li , Lijun Zhou , Sixu Yan , Bencheng Liao , Tianyi Yan , Kaixin Xiong , Long Chen , Hongwei Xie , Bing Wang , Guang Chen , Hangjun Ye , Wenyu Liu , Haiyang Sun , Xinggang Wang

Leveraging multiple sensors is crucial for robust semantic perception in autonomous driving, as each sensor type has complementary strengths and weaknesses. However, existing sensor fusion methods often treat sensors uniformly across all…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Tim Broedermann , Christos Sakaridis , Yuqian Fu , Luc Van Gool

This short paper presents a preliminary analysis of three popular Visual Question Answering (VQA) models, namely ViLBERT, ViLT, and LXMERT, in the context of answering questions relating to driving scenarios. The performance of these models…

计算机视觉与模式识别 · 计算机科学 2023-07-31 Kaavya Rekanar , Ciarán Eising , Ganesh Sistu , Martin Hayes

Recent research on Large Language Models for autonomous driving shows promise in planning and control. However, high computational demands and hallucinations still challenge accurate trajectory prediction and control signal generation.…

机器人学 · 计算机科学 2024-10-03 Ziang Guo , Zakhar Yagudin , Artem Lykov , Mikhail Konenkov , Dzmitry Tsetserukou

Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily…

计算与语言 · 计算机科学 2025-03-11 Yanling Wang , Yihan Zhao , Xiaodong Chen , Shasha Guo , Lixin Liu , Haoyang Li , Yong Xiao , Jing Zhang , Qi Li , Ke Xu

Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to maintain the same level of effectiveness in more specialized…

计算与语言 · 计算机科学 2024-04-26 Cuong Nhat Ha , Shima Asaadi , Sanjeev Kumar Karn , Oladimeji Farri , Tobias Heimann , Thomas Runkler

Vision Language Models (VLMs) bridge visual perception and linguistic reasoning. In Autonomous Driving (AD), this synergy has enabled Vision Language Action (VLA) models, which translate high-level multimodal understanding into driving…

机器人学 · 计算机科学 2026-03-11 Yuan Gao , Dengyuan Hua , Mattia Piccinini , Finn Rasmus Schäfer , Korbinian Moller , Lin Li , Johannes Betz

While multi-modal models have successfully integrated information from image, video, and audio modalities, integrating graph modality into large language models (LLMs) remains unexplored. This discrepancy largely stems from the inherent…

计算与语言 · 计算机科学 2023-10-13 Yuanchun Shen , Ruotong Liao , Zhen Han , Yunpu Ma , Volker Tresp

In this work, we study how vision-language models (VLMs) can be utilized to enhance the safety for the autonomous driving system, including perception, situational understanding, and path planning. However, existing research has largely…

人工智能 · 计算机科学 2025-07-30 Hao Ye , Mengshi Qi , Zhaohong Liu , Liang Liu , Huadong Ma

Traditional approaches to safety event analysis in autonomous systems have relied on complex machine learning models and extensive datasets for high accuracy and reliability. However, the advent of Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Mohammad Abu Tami , Huthaifa I. Ashqar , Mohammed Elhenawy

Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in environmental conditions and sensor inputs. As a result, they struggle to adaptively weight…

Cyclists often encounter safety-critical situations in urban traffic, highlighting the need for assistive systems that support safe and informed decision-making. Recently, vision-language models (VLMs) have demonstrated strong performance…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Krishna Kanth Nakka , Vedasri Nakka