中文
相关论文

相关论文: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 篇论文

Deploying autonomous edge robotics in dynamic military environments is constrained by both scarce domain-specific training data and the computational limits of edge hardware. This paper introduces a hierarchical, zero-shot framework that…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Jesse Barkley , Abraham George , Amir Barati Farimani

End-to-end learning has emerged as a major paradigm for developing autonomous systems. Unfortunately, with its performance and convenience comes an even greater challenge of safety assurance. A key factor of this challenge is the absence of…

机器学习 · 计算机科学 2024-06-21 Zhenjiang Mao , Carson Sobolewski , Ivan Ruchkin

Pedestrian-vehicle incidents remain a critical urban safety challenge, with pedestrians accounting for over 20% of global traffic fatalities. Although existing video-based systems can detect when incidents occur, they provide little insight…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Hao Zhen , Yunxiang Yang , Jidong J. Yang

Traffic safety analysis requires complex video understanding to capture fine-grained behavioral patterns and generate comprehensive descriptions for accident prevention. In this work, we present a unique dual-model framework that…

Video understanding is an important problem in computer vision. Currently, the well-studied task in this research is human action recognition, where the clips are manually trimmed from the long videos, and a single class of human action is…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Yi Liu , Xuan Zhang , Ying Li , Guixin Liang , Yabing Jiang , Lixia Qiu , Haiping Tang , Fei Xie , Wei Yao , Yi Dai , Yu Qiao , Yali Wang

Pedestrian occlusion is challenging for autonomous vehicles (AVs) at midblock locations on multilane roadways because an AV cannot detect crossing pedestrians that are fully occluded by downstream vehicles in adjacent lanes. This paper…

机器人学 · 计算机科学 2023-03-24 Fengjiao Zou , Hsien-Wen Deng , Tsing-Un Iunn , Jennifer Harper Ogle , Weimin Jin

The escalating intensity and frequency of wildfires demand innovative computational methods for rapid and accurate property damage assessment. Traditional methods are often time-consuming, while modern computer vision approaches typically…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Miguel Esparza , Archit Gupta , Kai Yin , Yiming Xiao , Ali Mostafavi

Generating intelligent robot behavior in contact-rich settings is a research problem where zeroth-order methods currently prevail. Developing methods that make use of first/second order information about rigid-body dynamics in the presence…

机器人学 · 计算机科学 2026-05-26 Onur Beker , Andreas René Geist , Anselm Paulus , Georg Martius

Estimating and understanding the current scene is an inevitable capability of automated vehicles. Usually, maps are used as prior for interpreting sensor measurements in order to drive safely and comfortably. Only few approaches take into…

计算机视觉与模式识别 · 计算机科学 2019-08-08 Annika Meyer , Jonas Walter , Martin Lauer , Christoph Stiller

Single-vehicle Vision-Language Models (VLMs) are fundamentally constrained by sensor occlusions. While Vehicle-to-Everything (V2X) systems mitigate this, current benchmarks lack the cooperative reasoning required for resolving ambiguities…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Kevin Richard , Alphin Varghese , Colin Pham , David Oh , Srijan Das

We introduce ACCIDENT, a benchmark dataset for traffic accident detection in CCTV footage, designed to evaluate models in supervised (IID and OOD) and zero-shot settings, reflecting both data-rich and data-scarce scenarios. The benchmark…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Lukas Picek , Michal Čermák , Marek Hanzl , Vojtěch Čermák

Pipeline parallelism is one of the key components for large-scale distributed training, yet its efficiency suffers from pipeline bubbles which were deemed inevitable. In this work, we introduce a scheduling strategy that, to our knowledge,…

分布式、并行与集群计算 · 计算机科学 2024-01-22 Penghui Qi , Xinyi Wan , Guangxing Huang , Min Lin

Effective risk monitoring in dynamic environments such as disaster zones requires an adaptive exploration strategy to detect hidden threats. We propose a bi-level unmanned aerial vehicle (UAV) monitoring strategy that efficiently integrates…

最优化与控制 · 数学 2026-01-22 Jimin Choi , Grant Stagg , Cameron K. Peterson , Max Z. Li

Addressing hard cases in autonomous driving, such as anomalous road users, extreme weather conditions, and complex traffic interactions, presents significant challenges. To ensure safety, it is crucial to detect and manage these scenarios…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Yi Yang , Qingwen Zhang , Kei Ikemura , Nazre Batool , John Folkesson

The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-collisions, scenarios that are brief, rare, and difficult for generic vision models to capture.…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Mohammad Qazim Bhat , Yufan Huang , Niket Agarwal , Hao Wang , Michael Woods , John Kenyon , Tsung-Yi Lin , Xiaodong Yang , Ming-Yu Liu , Kevin Xie

Recent developments in vision language models (VLM) have shown great potential for diverse applications related to image understanding. In this study, we have explored state-of-the-art VLM models for vision-based transportation engineering…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Sanjita Prajapati , Tanu Singh , Chinmay Hegde , Pranamesh Chakraborty

In this paper, we introduce VisioPath, a novel framework combining vision-language models (VLMs) with model predictive control (MPC) to enable safe autonomous driving in dynamic traffic environments. The proposed approach leverages a…

系统与控制 · 电气工程与系统科学 2025-07-10 Shanting Wang , Panagiotis Typaldos , Chenjun Li , Andreas A. Malikopoulos

Zero-shot vision-and-language navigation (VLN) has gained significant attention due to its minimal data collection costs and inherent generalization. This paradigm is typically driven by the integration of pre-trained Vision-Language Models…

机器人学 · 计算机科学 2026-05-15 Ziyi Xia , Chaoran Xiong , Litao Wei , Xinhao Hu , Ling Pei

Traffic prediction has long been a focal and pivotal area in research, witnessing both significant strides from city-level to road-level predictions in recent years. With the advancement of Vehicle-to-Everything (V2X) technologies,…

机器学习 · 计算机科学 2025-06-17 Shuhao Li , Yue Cui , Jingyi Xu , Libin Li , Lingkai Meng , Weidong Yang , Fan Zhang , Xiaofang Zhou

Multi-step zoom-in pipelines are widely used for GUI grounding, yet the intermediate predictions they produce are typically discarded after coordinate remapping. We observe that these intermediate outputs contain a useful confidence signal…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Keon Kim , Krish Chelikavada