中文
相关论文

相关论文: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 篇论文

Foundation models, especially vision-language models (VLMs), offer compelling zero-shot object detection for applications like autonomous driving, a domain where manual labelling is prohibitively expensive. However, their detection latency…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Uday Bhaskar , Rishabh Bhattacharya , Avinash Patel , Sarthak Khoche , Praveen Anil Kulkarni , Naresh Manwani

3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods predominantly operate in the 2D domain, where viewpoint variations are mistaken for…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Wei Zhang , Songhua Li , Yihang Wu , Qiang Li , Qi Wang

Obstacle avoidance is essential for ensuring the safety of autonomous vehicles. Accurate perception and motion planning are crucial to enabling vehicles to navigate complex environments while avoiding collisions. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Van-Hoang-Anh Phan , Chi-Tam Nguyen , Doan-Trung Au , Thanh-Danh Phan , Minh-Thien Duong , My-Ha Le

Risk quantification is a critical component of safe autonomous driving, however, constrained by the limited perception range and occlusion of single-vehicle systems in complex and dense scenarios. Vehicle-to-everything (V2X) paradigm has…

机器人学 · 计算机科学 2025-06-23 Mingyue Lei , Zewei Zhou , Hongchen Li , Jia Hu , Jiaqi Ma

Autonomous Vehicles (AVs) are often tested in simulation to estimate the probability they will violate safety specifications. Two common issues arise when using existing techniques to produce this estimation: If violations occur rarely,…

机器人学 · 计算机科学 2024-07-25 Craig Innes , Subramanian Ramamoorthy

Injecting world knowledge into pretrained multimodal large language models (MLLMs) is essential for domain-specific applications. Task-specific fine-tuning achieves this by tailoring MLLMs to high-quality in-domain data but encounters…

多媒体 · 计算机科学 2026-03-31 Xiao An , Jiaxing Sun , Ting Hu , Wei He

Indoor environments lack the spatial intelligence infrastructure that GPS provides outdoors; first responders arriving at unfamiliar buildings typically have no machine-readable map of safety equipment. Prior work on 3D semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Alexander Nikitas Dimopoulos , Joseph Grasso , John Beltz

Despite significant results achieved by Contrastive Language-Image Pretraining (CLIP) in zero-shot image recognition, limited effort has been made exploring its potential for zero-shot video recognition. This paper presents Open-VCLIP++, a…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Zuxuan Wu , Zejia Weng , Wujian Peng , Xitong Yang , Ang Li , Larry S. Davis , Yu-Gang Jiang

Visual Place Recognition (VPR) enables robots and autonomous vehicles to identify previously visited locations by matching current observations against a database of known places. However, VPR systems face significant challenges when…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Emily Miller , Michael Milford , Muhammad Burhan Hafez , SD Ramchurn , Shoaib Ehsan

In this paper, we focus on the performance analysis of a semi-persistent scheduling scheme for vehicular safety communications, motivated by the Mode 4 medium access control protocol in 3GPP Release 14 for Cellular-V2X. An analytical model…

网络与互联网体系结构 · 计算机科学 2018-08-30 Xu Wang , Randall A. Berry , Ivan Vukovic , Jayanthi Rao

Autonomous driving is a complex and challenging task that aims at safe motion planning through scene understanding and reasoning. While vision-only autonomous driving methods have recently achieved notable performance, through enhanced…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Chenbin Pan , Burhaneddin Yaman , Tommaso Nesti , Abhirup Mallik , Alessandro G Allievi , Senem Velipasalar , Liu Ren

Variable Speed Limit (VSL) control has been one of the most popular techniques with the potential of smoothing traffic flow, maximizing throughput at bottlenecks, and improving mobility and safety. Despite the substantial research efforts…

系统与控制 · 电气工程与系统科学 2025-05-12 Tianchen Yuan , Faisal Alasiri , Petros A. Ioannou

Adjusting rifle sights, a process commonly called "zeroing," requires shooters to identify and differentiate bullet holes from multiple firing iterations. Traditionally, this process demands physical inspection, introducing delays due to…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Robert M. Belcher , Brendan C. Degryse , Leonard R. Kosta , Christopher J. Lowrance

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts…

计算机视觉与模式识别 · 计算机科学 2025-05-02 Jannik Lübberstedt , Esteban Rivera , Nico Uhlemann , Markus Lienkamp

Adopting contrastive image-text pretrained models like CLIP towards video classification has gained attention due to its cost-effectiveness and competitive performance. However, recent works in this area face a trade-off. Finetuning the…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Syed Talal Wasim , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan , Mubarak Shah

Accurate prediction of traffic crash risks for individual vehicles is essential for enhancing vehicle safety. While significant attention has been given to traffic crash risk prediction, existing studies face two main challenges: First, due…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Kequan Chen , Pan Liu , Yuxuan Wang , David Z. W. Wang , Yifan Dai , Zhibin Li

Video prediction models based on convolutional networks, recurrent networks, and their combinations often result in blurry predictions. We identify an important contributing factor for imprecise predictions that has not been studied…

计算机视觉与模式识别 · 计算机科学 2018-09-11 Wonmin Byeon , Qin Wang , Rupesh Kumar Srivastava , Petros Koumoutsakos

We present AutoSiMP, an autonomous pipeline that transforms a natural-language structural problem description into a validated, binary topology without manual configuration. The pipeline comprises five modules: (1) an LLM-based configurator…

计算工程、金融与科学 · 计算机科学 2026-03-31 Shaoliang Yang , Jun Wang , Yunsheng Wang

An advanced Volume of Fluid (VOF) method is presented that enables performant three-dimensional Direct Numerical Simulations (DNS) of the interaction of two immiscible fluids in a gaseous environment with large topology changes, e.g.,…

流体动力学 · 物理学 2023-11-08 Johanna Potyka , Kathrin Schulte

3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Nikita Drozdov , Andrey Lemeshko , Nikita Gavrilov , Anton Konushin , Danila Rukhovich , Maksim Kolodiazhnyi