中文
相关论文

相关论文: ScVLM: Enhancing Vision-Language Model for Safety-…

200 篇论文

The emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multimodal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Yongting Zhang , Lu Chen , Guodong Zheng , Yifeng Gao , Rui Zheng , Jinlan Fu , Zhenfei Yin , Senjie Jin , Yu Qiao , Xuanjing Huang , Feng Zhao , Tao Gui , Jing Shao

Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Haoxiang Gao , Li Zhang , Yu Zhao , Zhou Yang , Jinghan Cao

Safety testing serves as the fundamental pillar for the development of autonomous driving systems (ADSs). To ensure the safety of ADSs, it is paramount to generate a diverse range of safety-critical test scenarios. While existing ADS…

软件工程 · 计算机科学 2025-01-03 Haoxiang Tian , Xingshuo Han , Yuan Zhou , Guoquan Wu , An Guo , Mingfei Cheng , Shuo Li , Jun Wei , Tianwei Zhang

This research showcases the innovative integration of Large Language Models into machine learning workflows for traffic incident management, focusing on the classification of incident severity using accident reports. By leveraging features…

机器学习 · 计算机科学 2024-05-01 Artur Grigorev , Khaled Saleh , Yuming Ou , Adriana-Simona Mihaita

In this study, observations of the Vocational Education and Training (VET) in mechanical engineering companies are carried out. A Learning Management System (LMS) had been developed for the assistance in solving typical task structures,…

计算机与社会 · 计算机科学 2019-01-15 Adrian Wilke , Johannes Magenheim

Autonomous Vehicles (AVs) collect and pseudo-label terabytes of multi-modal data localized to HD maps during normal fleet testing. However, identifying interesting and safety-critical scenarios from uncurated driving logs remains a…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Cainan Davidson , Deva Ramanan , Neehar Peri

Prior studies on Video Anomaly Detection (VAD) mainly focus on detecting whether each video frame is abnormal or not in the video, which largely ignore the structured video semantic information (i.e., what, when, and where does the abnormal…

计算机视觉与模式识别 · 计算机科学 2025-02-27 Junxiao Ma , Jingjing Wang , Jiamin Luo , Peiying Yu , Guodong Zhou

Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and grounding. Existing methods, such as proxy encoding and geometry…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Tianren Ma , Lingxi Xie , Yunjie Tian , Boyu Yang , Qixiang Ye

Recent years have witnessed a significant increase in the performance of Vision and Language tasks. Foundational Vision-Language Models (VLMs), such as CLIP, have been leveraged in multiple settings and demonstrated remarkable performance…

计算机视觉与模式识别 · 计算机科学 2024-03-04 Santiago Castro , Amir Ziai , Avneesh Saluja , Zhuoning Yuan , Rada Mihalcea

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequential decision-making…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Weizhen Wang , Chenda Duan , Zhenghao Peng , Yuxin Liu , Bolei Zhou

Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations. We hypothesize that this issue…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Shengguang Wu , Fan-Yun Sun , Kaiyue Wen , Nick Haber

Accurate prediction of traffic crash risks for individual vehicles is essential for enhancing vehicle safety. While significant attention has been given to traffic crash risk prediction, existing studies face two main challenges: First, due…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Kequan Chen , Pan Liu , Yuxuan Wang , David Z. W. Wang , Yifan Dai , Zhibin Li

The safe deployment of autonomous driving systems (ADSs) relies on comprehensive testing and evaluation. However, safety-critical scenarios that can effectively expose system vulnerabilities are extremely sparse in the real world. Existing…

机器人学 · 计算机科学 2025-12-03 Xinzheng Wu , Junyi Chen , Naiting Zhong , Yong Shen

Vision-Language Models (VLMs) excel in diverse multimodal tasks. However, user requirements vary across scenarios, which can be categorized into fast response, high-quality output, and low energy consumption. Relying solely on large models…

机器学习 · 计算机科学 2025-11-03 Xin Tang , Youfang Han , Fangfei Gou , Wei Zhao , Xin Meng , Yang Yu , Jinguo Zhang , Yuanchun Shi , Yuntao Wang , Tengxiang Zhang

Road crashes claim over 1.3 million lives annually worldwide and incur global economic losses exceeding \$1.8 trillion. Such profound societal and financial impacts underscore the urgent need for road safety research that uncovers crash…

计算与语言 · 计算机科学 2025-05-14 Hao Zhen , Jidong J. Yang

Sparse Autoencoders (SAEs) have recently gained attention as a means to improve the interpretability and steerability of Large Language Models (LLMs), both of which are essential for AI safety. In this work, we extend the application of…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Mateusz Pach , Shyamgopal Karthik , Quentin Bouniot , Serge Belongie , Zeynep Akata

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

人工智能 · 计算机科学 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns. This issue arises from the lack of comprehensive benchmarks that…

Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses under near-identical counterfactual scenes. We present…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Xingcheng Zhou , Hao Guo , Rui Song , Walter Zimmer , Mingyu Liu , André Schamschurko , Hu Cao , Alois Knoll

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations…

计算机视觉与模式识别 · 计算机科学 2022-05-17 Zirui Wang , Jiahui Yu , Adams Wei Yu , Zihang Dai , Yulia Tsvetkov , Yuan Cao