English
Related papers

Related papers: ScVLM: Enhancing Vision-Language Model for Safety-…

200 papers

Data for training learning-enabled self-driving cars in the physical world are typically collected in a safe, normal environment. Such data distribution often engenders a strong bias towards safe driving, making self-driving cars unprepared…

End-to-end autonomous driving demonstrates strong planning capabilities with large-scale data but still struggles in complex, rare scenarios due to limited commonsense. In contrast, Large Vision-Language Models (LVLMs) excel in scene…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Bo Jiang , Shaoyu Chen , Bencheng Liao , Xingyu Zhang , Wei Yin , Qian Zhang , Chang Huang , Wenyu Liu , Xinggang Wang

Modern cities experience heavy traffic flows and congestions regularly across space and time. Monitoring traffic situations becomes an important challenge for the Traffic Control and Surveillance Systems (TCSS). In advanced TCSS, it is…

Machine Learning · Computer Science 2015-12-29 Li-Li Wang , Henry Y. T. Ngan , Nelson H. C. Yung

In the field of Class Incremental Object Detection (CIOD), creating models that can continuously learn like humans is a major challenge. Pseudo-labeling methods, although initially powerful, struggle with multi-scenario incremental learning…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Junsu Kim , Yunhoe Ku , Jihyeon Kim , Junuk Cha , Seungryul Baek

Vision-Language-Action (VLA)-based driving systems represent a significant paradigm shift in autonomous driving since, by combining traffic scene understanding, linguistic interpretation, and action generation, these systems enable more…

Robotics · Computer Science 2026-03-19 Gerhard Yu , Fuyuki Ishikawa , Oluwafemi Odu , Alvine Boaye Belle

This study examines the feasibility of applying large language models (LLMs) for forecasting the impact of traffic incidents on the traffic flow. The use of LLMs for this task has several advantages over existing machine learning-based…

Artificial Intelligence · Computer Science 2025-07-08 George Jagadeesh , Srikrishna Iyer , Michal Polanowski , Kai Xin Thia

Assessing scenario coverage is crucial for evaluating the robustness of autonomous agents, yet existing methods rely on expensive human annotations or computationally intensive Large Vision-Language Models (LVLMs). These approaches are…

Robotics · Computer Science 2025-10-30 Anil Yildiz , Sarah M. Thornton , Carl Hildebrandt , Sreeja Roy-Singh , Mykel J. Kochenderfer

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To address this problem,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Hsu-kuang Chiu , Ryo Hachiuma , Chien-Yi Wang , Stephen F. Smith , Yu-Chiang Frank Wang , Min-Hung Chen

Traditional autonomous driving systems often struggle to connect high-level reasoning with low-level control, leading to suboptimal and sometimes unsafe behaviors. Recent advances in multimodal large language models (MLLMs), which process…

Robotics · Computer Science 2025-06-09 Jiawei Zhang , Xuan Yang , Taiqi Wang , Yu Yao , Aleksandr Petiushko , Bo Li

Traffic safety analysis requires complex video understanding to capture fine-grained behavioral patterns and generate comprehensive descriptions for accident prevention. In this work, we present a unique dual-model framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Blessing Agyei Kyem , Neema Jakisa Owor , Andrews Danyo , Joshua Kofi Asamoah , Eugene Denteh , Tanner Muturi , Anthony Dontoh , Yaw Adu-Gyamfi , Armstrong Aboah

Vision language models (VLM) demonstrate sophisticated multimodal reasoning yet are prone to hallucination when confronted with knowledge conflicts, impeding their deployment in information-sensitive contexts. While existing research…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Peter Carragher , Nikitha Rao , Abhinand Jha , R Raghav , Kathleen M. Carley

Vision-Language Models (VLMs) have recently emerged as a promising paradigm in autonomous driving (AD). However, current performance evaluation protocols for VLM-based AD systems (ADVLMs) are predominantly confined to open-loop settings…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Tianyuan Zhang , Ting Jin , Lu Wang , Jiangfan Liu , Siyuan Liang , Mingchuan Zhang , Aishan Liu , Xianglong Liu

Vision-Language Models (VLMs) have advanced multi-modal tasks like image captioning, visual question answering, and reasoning. However, they often generate hallucinated outputs inconsistent with the visual context or prompt, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Shawn Li , Jiashu Qu , Yuxiao Zhou , Yuehan Qin , Tiankai Yang , Yue Zhao

Video captioning is a critical task in the field of multimodal machine learning, aiming to generate descriptive and coherent textual narratives for video content. While large vision-language models (LVLMs) have shown significant progress,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Ji-jun Park , Soo-joon Choi

The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long tail scenarios. However, these models…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Nikos Theodoridis , Reenu Mohandas , Ganesh Sistu , Anthony Scanlan , Ciarán Eising , Tim Brophy

Today's advanced driver assistance systems (ADAS), like adaptive cruise control or rear collision warning, are finding broader adoption across vehicle classes. Integrating such advanced, multimodal Large Language Models (LLMs) on board a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Malsha Ashani Mahawatta Dona , Beatriz Cabrero-Daniel , Yinan Yu , Christian Berger

Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accurate visual understanding. To address this issue, we propose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Woohyeon Park , Woojin Kim , Jaeik Kim , Jaeyoung Do

The unpredictable nature of outdoor settings introduces numerous safety concerns, making hazard detection crucial for safe navigation. This paper introduces a novel system for sidewalk safety navigation utilizing a hybrid approach that…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Edgar Guzman , Robert D. Howe

Vision-language models (VLMs) have significantly improved the generalization capabilities of robotic manipulation. However, VLM-based systems often suffer from a lack of robustness, leading to unpredictable errors, particularly in scenarios…

Robotics · Computer Science 2026-03-17 Yayun He , Zuheng Kang , Botao Zhao , Zhouyin Wu , Junqing Peng , Jianzong Wang

Recent research on Large Language Models for autonomous driving shows promise in planning and control. However, high computational demands and hallucinations still challenge accurate trajectory prediction and control signal generation.…

Robotics · Computer Science 2024-10-03 Ziang Guo , Zakhar Yagudin , Artem Lykov , Mikhail Konenkov , Dzmitry Tsetserukou