中文
相关论文

相关论文: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 篇论文

In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing with the representation of state-action mapping in the…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Ziang Guo , Konstantin Gubernatorov , Selamawit Asfaw , Zakhar Yagudin , Dzmitry Tsetserukou

Predicting on-road abnormalities such as road accidents or traffic violations is a challenging task in traffic surveillance. If such predictions can be done in advance, many damages can be controlled. Here in our wok, we tried to formulate…

计算机视觉与模式识别 · 计算机科学 2021-01-22 Deesha Chavan , Dev Saad , Debarati B. Chakraborty

Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These models can plan a sequence of motions by evaluating the current view at each time step…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Wanrong Zheng , Yunhao Ge , Laurent Itti

This paper proposes a centralized multi-vehicle coordination scheme serving unsignalized intersections. The whole process consists of three stages: a) target velocity optimization: formulate the collision-free vehicle coordination as a…

系统与控制 · 电气工程与系统科学 2020-09-01 Qiang Ge , Qi Sun , Zhen Wang , Shengbo Eben Li , Ziqing Gu , Sifa Zheng

Reliable pipeline inspection is critical to safe energy transportation, but is constrained by long distances, complex terrain, and risks to human inspectors. Unmanned aerial vehicles provide a flexible sensing platform, yet reliable…

机器人学 · 计算机科学 2026-04-22 Wen Li , Hui Wang , Jinya Su , Cunjia Liu , Wen-Hua Chen , Shihua Li

The recent emergence of multimodal large language models (LLMs) has introduced new opportunities for improving visual hazard recognition on construction sites. Unlike traditional computer vision models that rely on domain-specific training…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Nishi Chaudhary , S M Jamil Uddin , Sathvik Sharath Chandra , Anto Ovid , Alex Albert

In this work, a multi-stage Machine Learning (ML) pipeline is proposed for pipe leakage detection in an industrial environment. As opposed to other industrial and urban environments, the environment under study includes many interfering…

机器学习 · 计算机科学 2022-05-06 Ibrahim Shaer , Abdallah Shami

Traditional approaches to safety event analysis in autonomous systems have relied on complex machine learning models and extensive datasets for high accuracy and reliability. However, the advent of Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Mohammad Abu Tami , Huthaifa I. Ashqar , Mohammed Elhenawy

This paper presents a general end-to-end framework for constructing robust and reliable layered safety filters that can be leveraged to perform dynamic collision avoidance over a broad range of applications using only local perception data.…

机器人学 · 计算机科学 2026-03-03 Erina Yamaguchi , Ryan M. Bena , Gilbert Bahati , Aaron D. Ames

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient solutions for…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Ilhan Skender , Kailin Tong , Selim Solmaz , Daniel Watzenig

Construction safety inspection remains mostly manual, and automated approaches still rely on task-specific datasets that are hard to maintain in fast-changing construction environments due to frequent retraining. Meanwhile, field inspection…

机器人学 · 计算机科学 2025-12-17 Hossein Naderi , Alireza Shojaei , Philip Agee , Kereshmeh Afsari , Abiola Akanmu

The recent large-scale vision-language pre-training (VLP) of dual-stream architectures (e.g., CLIP) with a tremendous amount of image-text pair data, has shown its superiority on various multimodal alignment tasks. Despite its success, the…

计算与语言 · 计算机科学 2022-03-31 Wenliang Dai , Lu Hou , Lifeng Shang , Xin Jiang , Qun Liu , Pascale Fung

The paper presents a vision-based obstacle avoidance strategy for lightweight self-driving cars that can be run on a CPU-only device using a single RGB-D camera. The method consists of two steps: visual perception and path planning. The…

机器人学 · 计算机科学 2024-08-22 Zhihao Lin , Zhen Tian , Qi Zhang , Hanyang Zhuang , Jianglin Lan

Localization is a critical technology in autonomous driving, encompassing both topological localization, which identifies the most similar map keyframe to the current observation, and metric localization, which provides precise spatial…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Ze Huang , Zhongyang Xiao , Mingliang Song , Longan Yang , Hongyuan Yuan , Li Sun

Anticipating the motion of neighboring vehicles is crucial for autonomous driving, especially on congested highways where even slight motion variations can result in catastrophic collisions. An accurate prediction of a future trajectory…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Fuad Hasan , Hailong Huang

Traffic scene understanding is essential for intelligent transportation systems and autonomous driving, ensuring safe and efficient vehicle operation. While recent advancements in VLMs have shown promise for holistic scene understanding,…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Qingyao Xu , Siheng Chen , Guang Chen , Yanfeng Wang , Ya Zhang

Detecting anomalous hazards in visual data, particularly in video streams, is a critical challenge in autonomous driving. Existing models often struggle with unpredictable, out-of-label hazards due to their reliance on predefined object…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Shashank Shriram , Srinivasa Perisetla , Aryan Keskar , Harsha Krishnaswamy , Tonko Emil Westerhof Bossen , Andreas Møgelmose , Ross Greer

Estimating the current scene and understanding the potential maneuvers are essential capabilities of automated vehicles. Most approaches rely heavily on the correctness of maps, but neglect the possibility of outdated information. We…

机器人学 · 计算机科学 2020-07-15 Annika Meyer , Jonas Walter , Martin Lauer

Motorcycle accidents pose significant risks, particularly when riders and passengers do not wear helmets. This study evaluates the efficacy of an advanced vision-language foundation model, OWLv2, in detecting and classifying various…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Lucas Choi , Ross Greer

We present a new, for plasma physics, highly efficient multilevel Monte Carlo numerical method for simulating Coulomb collisions. The method separates and optimally minimizes the finite-timestep and finite-sampling errors inherent in the…

等离子体物理 · 物理学 2015-08-12 M. S. Rosin , L. F. Ricketson , A. M. Dimits , R. E. Caflisch , B. I. Cohen