中文
相关论文

相关论文: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 篇论文

We describe a zero-shot pipeline developed for the ACCIDENT @ CVPR 2026 challenge. The challenge requires predicting when, where, and what type of traffic accident occurs in surveillance video, without labeled real-world training data. Our…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Amey Thakur , Sarvesh Talele

Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate joint localization in time, space, and collision type is required. We present a…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Jiantang Huang

We present the 1st-place solution of OpenLane Topology in Autonomous Driving Challenge. Considering that topology reasoning is based on centerline detection and traffic element detection, we develop a multi-stage framework for high…

计算机视觉与模式识别 · 计算机科学 2023-06-19 Dongming Wu , Fan Jia , Jiahao Chang , Zhuoling Li , Jianjian Sun , Chunrui Han , Shuailin Li , Yingfei Liu , Zheng Ge , Tiancai Wang

Terrain traversability estimation is crucial for autonomous robots, especially in unstructured environments where visual cues and reasoning play a key role. While vision-language models (VLMs) offer potential for zero-shot estimation, the…

机器人学 · 计算机科学 2025-08-05 Ida Germann , Mark O. Mints , Peer Neubert

Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Peter Robicheaux , Matvei Popov , Anish Madan , Isaac Robinson , Joseph Nelson , Deva Ramanan , Neehar Peri

Traditional Automatic License Plate Recognition (ALPR) systems employ multi-stage pipelines consisting of object detection networks followed by separate Optical Character Recognition (OCR) modules, introducing compounding errors, increased…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Karthik Sivakoti

Autonomous driving increasingly relies on Visual Question Answering (VQA) to enable vehicles to understand complex surroundings by analyzing visual inputs and textual queries. Currently, a paramount concern for VQA in this domain is the…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Yuliang Cai , Dongqiangzi Ye , Zitian Chen , Chongruo Wu

Infrastructure in smart cities is increasingly monitored by networks of closed circuit television (CCTV) cameras. Roads, bridges and tunnels develop cracks, potholes, and fluid leaks that threaten public safety and require timely repair.…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Ibrahim Sheikh Mohamed , Abdullah Yahya Abdullah Omaisan

Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Rui Gan , Junyi Ma , Pei Li , Xingyou Yang , Kai Chen , Sikai Chen , Bin Ran

This paper presents our submission to the COOOL competition, a novel benchmark for detecting and classifying out-of-label hazards in autonomous driving. Our approach integrates diverse methods across three core tasks: (i) driver reaction…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Lukas Picek , Vojtěch Čermák , Marek Hanzl

We propose the zero-shot Vision-and-Language Navigation with Collision Mitigation (VLN-CM), which takes these considerations. VLN-CM is composed of four modules and predicts the direction and distance of the next movement at each step. We…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Seongjun Jeong , Gi-Cheon Kang , Joochan Kim , Byoung-Tak Zhang

Autonomous inspection of underground infrastructure, such as sewer and culvert systems, is critical to public safety and urban sustainability. Although robotic platforms equipped with visual sensors can efficiently detect structural…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Johny J. Lopez , Md Meftahul Ferdaus , Mahdi Abdelguerfi

Construction remains the deadliest industry sector in the United States, with 1,055 fatal worker injuries recorded in 2023, and the majority preventable. Existing monitoring approaches are expensive, require real-time human operators, or…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Ananth Sriram , Neel Mokaria , Rajveer Singh

Autonomous driving systems often infer pedestrian yielding behavior from geometric and kinematic cues alone, limiting their ability to reason about visual scene context and age-dependent behavioral variability. This limitation can produce…

系统与控制 · 电气工程与系统科学 2026-04-28 Qingwen Pu , Kun Xie , Yuxiang Liu

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Tomaso Trinci , Henrique Piñeiro Monteagudo , Leonardo Taccari

With the rapid progress of foundation models and robotics, vision-language navigation (VLN) has emerged as a key task for embodied agents with broad practical applications. We address VLN in continuous environments, a particularly…

机器人学 · 计算机科学 2025-09-26 Boqi Li , Siyuan Li , Weiyi Wang , Anran Li , Zhong Cao , Henry X. Liu

Bridge infrastructure inspection is a critical but labor-intensive task requiring expert assessment of structural damage such as rebar exposure, cracking, and corrosion. This paper presents a comprehensive study of quantized Vision-Language…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Takato Yasuno

Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, i.e.,…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Dongming Wu , Jiahao Chang , Fan Jia , Yingfei Liu , Tiancai Wang , Jianbing Shen

Crash diagrams are essential tools in transportation safety analysis, yet their manual preparation remains time-consuming and prone to human variability. This study investigates the use of Vision-Language Models (VLMs) to automate crash…

人机交互 · 计算机科学 2026-04-20 Xiao Lu , Hao Zhen , Jidong J. Yang

This paper introduces an efficient Vision-Language Model (VLM) pipeline specifically optimized for deployment on embedded devices, such as those used in robotics and autonomous driving. The pipeline significantly reduces the computational…

机器学习 · 计算机科学 2025-11-04 Jin Huang , Yuchao Jin , Le An , Josh Park
‹ 上一页 1 2 3 10 下一页 ›