English
Related papers

Related papers: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 papers

In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing with the representation of state-action mapping in the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Ziang Guo , Konstantin Gubernatorov , Selamawit Asfaw , Zakhar Yagudin , Dzmitry Tsetserukou

Predicting on-road abnormalities such as road accidents or traffic violations is a challenging task in traffic surveillance. If such predictions can be done in advance, many damages can be controlled. Here in our wok, we tried to formulate…

Computer Vision and Pattern Recognition · Computer Science 2021-01-22 Deesha Chavan , Dev Saad , Debarati B. Chakraborty

Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These models can plan a sequence of motions by evaluating the current view at each time step…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wanrong Zheng , Yunhao Ge , Laurent Itti

This paper proposes a centralized multi-vehicle coordination scheme serving unsignalized intersections. The whole process consists of three stages: a) target velocity optimization: formulate the collision-free vehicle coordination as a…

Systems and Control · Electrical Eng. & Systems 2020-09-01 Qiang Ge , Qi Sun , Zhen Wang , Shengbo Eben Li , Ziqing Gu , Sifa Zheng

Reliable pipeline inspection is critical to safe energy transportation, but is constrained by long distances, complex terrain, and risks to human inspectors. Unmanned aerial vehicles provide a flexible sensing platform, yet reliable…

Robotics · Computer Science 2026-04-22 Wen Li , Hui Wang , Jinya Su , Cunjia Liu , Wen-Hua Chen , Shihua Li

The recent emergence of multimodal large language models (LLMs) has introduced new opportunities for improving visual hazard recognition on construction sites. Unlike traditional computer vision models that rely on domain-specific training…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Nishi Chaudhary , S M Jamil Uddin , Sathvik Sharath Chandra , Anto Ovid , Alex Albert

In this work, a multi-stage Machine Learning (ML) pipeline is proposed for pipe leakage detection in an industrial environment. As opposed to other industrial and urban environments, the environment under study includes many interfering…

Machine Learning · Computer Science 2022-05-06 Ibrahim Shaer , Abdallah Shami

Traditional approaches to safety event analysis in autonomous systems have relied on complex machine learning models and extensive datasets for high accuracy and reliability. However, the advent of Multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Mohammad Abu Tami , Huthaifa I. Ashqar , Mohammed Elhenawy

This paper presents a general end-to-end framework for constructing robust and reliable layered safety filters that can be leveraged to perform dynamic collision avoidance over a broad range of applications using only local perception data.…

Robotics · Computer Science 2026-03-03 Erina Yamaguchi , Ryan M. Bena , Gilbert Bahati , Aaron D. Ames

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient solutions for…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Ilhan Skender , Kailin Tong , Selim Solmaz , Daniel Watzenig

Construction safety inspection remains mostly manual, and automated approaches still rely on task-specific datasets that are hard to maintain in fast-changing construction environments due to frequent retraining. Meanwhile, field inspection…

Robotics · Computer Science 2025-12-17 Hossein Naderi , Alireza Shojaei , Philip Agee , Kereshmeh Afsari , Abiola Akanmu

The recent large-scale vision-language pre-training (VLP) of dual-stream architectures (e.g., CLIP) with a tremendous amount of image-text pair data, has shown its superiority on various multimodal alignment tasks. Despite its success, the…

Computation and Language · Computer Science 2022-03-31 Wenliang Dai , Lu Hou , Lifeng Shang , Xin Jiang , Qun Liu , Pascale Fung

The paper presents a vision-based obstacle avoidance strategy for lightweight self-driving cars that can be run on a CPU-only device using a single RGB-D camera. The method consists of two steps: visual perception and path planning. The…

Robotics · Computer Science 2024-08-22 Zhihao Lin , Zhen Tian , Qi Zhang , Hanyang Zhuang , Jianglin Lan

Localization is a critical technology in autonomous driving, encompassing both topological localization, which identifies the most similar map keyframe to the current observation, and metric localization, which provides precise spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Ze Huang , Zhongyang Xiao , Mingliang Song , Longan Yang , Hongyuan Yuan , Li Sun

Anticipating the motion of neighboring vehicles is crucial for autonomous driving, especially on congested highways where even slight motion variations can result in catastrophic collisions. An accurate prediction of a future trajectory…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Fuad Hasan , Hailong Huang

Traffic scene understanding is essential for intelligent transportation systems and autonomous driving, ensuring safe and efficient vehicle operation. While recent advancements in VLMs have shown promise for holistic scene understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Qingyao Xu , Siheng Chen , Guang Chen , Yanfeng Wang , Ya Zhang

Detecting anomalous hazards in visual data, particularly in video streams, is a critical challenge in autonomous driving. Existing models often struggle with unpredictable, out-of-label hazards due to their reliance on predefined object…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Shashank Shriram , Srinivasa Perisetla , Aryan Keskar , Harsha Krishnaswamy , Tonko Emil Westerhof Bossen , Andreas Møgelmose , Ross Greer

Estimating the current scene and understanding the potential maneuvers are essential capabilities of automated vehicles. Most approaches rely heavily on the correctness of maps, but neglect the possibility of outdated information. We…

Robotics · Computer Science 2020-07-15 Annika Meyer , Jonas Walter , Martin Lauer

Motorcycle accidents pose significant risks, particularly when riders and passengers do not wear helmets. This study evaluates the efficacy of an advanced vision-language foundation model, OWLv2, in detecting and classifying various…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Lucas Choi , Ross Greer

We present a new, for plasma physics, highly efficient multilevel Monte Carlo numerical method for simulating Coulomb collisions. The method separates and optimally minimizes the finite-timestep and finite-sampling errors inherent in the…

Plasma Physics · Physics 2015-08-12 M. S. Rosin , L. F. Ricketson , A. M. Dimits , R. E. Caflisch , B. I. Cohen