中文
相关论文

相关论文: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 篇论文

Relative vehicle positioning methods can contribute to safer and more efficient autonomous driving by enabling collision avoidance and platooning applications. For full automation, these applications require cm-level positioning accuracy…

信号处理 · 电气工程与系统科学 2023-11-06 Burak Soner , Merve Karakas , Utku Noyan , Furkan Sahbaz , Sinem Coleri

A pervasive intuition holds that vision-language models (VLMs) are most trustworthy when their attention maps look sharp: concentrated attention on the queried region should imply a confident, calibrated answer. We test this…

人工智能 · 计算机科学 2026-05-12 Logan Mann , Ajit Saravanan , Ishan Dave , Shikhar Shiromani , Saadullah Ismail , Yi Xia , Emily Huang

Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundation Models (VFMs) and Multimodal Large Language Models…

机器学习 · 计算机科学 2025-09-03 Jintao Cheng , Weibin Li , Jiehao Luo , Xiaoyu Tang , Zhijian He , Jin Wu , Yao Zou , Wei Zhang

Large-scale visual-language pre-trained models (VLPM) have proven their excellent performance in downstream object detection for natural scenes. However, zero-shot nuclei detection on H\&E images via VLPMs remains underexplored. The large…

计算机视觉与模式识别 · 计算机科学 2023-07-03 Yongjian Wu , Yang Zhou , Jiya Saiyin , Bingzheng Wei , Maode Lai , Jianzhong Shou , Yubo Fan , Yan Xu

Traditional license plate detection and recognition models are often trained on closed datasets, limiting their ability to handle the diverse license plate formats across different regions. The emergence of large-scale pre-trained models…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Haoxuan Ding , Qi Wang , Junyu Gao , Qiang Li

Trajectory prediction with uncertainty is a critical and challenging task for autonomous driving. Nowadays, we can easily access sensor data represented in multiple views. However, cross-view consistency has not been evaluated by the…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Zijian Song , Huikun Bi , Ruisi Zhang , Tianlu Mao , Zhaoqi Wang

To support large-scale model training, split learning (SL) enables multiple edge devices/servers to share the intensive training workload. However, most existing works on SL focus solely on two-tier model splitting. Moreover, while some…

网络与互联网体系结构 · 计算机科学 2025-09-19 Wei Wei , Zheng Lin , Tao Li , Xuanheng Li , Xianhao Chen

Event cameras, featuring high temporal resolution and high dynamic range, offer visual sensing capabilities comparable to conventional image sensors while capturing fast-moving objects and handling scenes with extreme lighting contrasts…

机器人学 · 计算机科学 2025-10-21 Ryota Soga , Masataka Kobayashi , Tsukasa Shimizu , Shintaro Shiba , Quan Kong , Shan Lu , Takaya Yamazato

Cutting-edge connected vehicle (CV) technologies have drawn much attention in recent years. The real-time traffic data captured by a CV can be shared with other CVs and data centers so as to open new possibilities for solving diverse…

计算机视觉与模式识别 · 计算机科学 2023-06-22 Shaocheng Jia , Wei Yao

The ability of autonomous vehicles to maintain an accurate trajectory within their road lane is crucial for safe operation. This requires detecting the road lines and estimating the car relative pose within its lane. Lateral lines are…

机器人学 · 计算机科学 2021-01-19 Paolo Cudrano , Simone Mentasti , Matteo Matteucci , Mattia Bersani , Stefano Arrigoni , Federico Cheli

Vision-language foundation models (VLMs) show promise for diverse imaging tasks but often underperform on medical benchmarks. Prior efforts to improve performance include model finetuning, which requires large domain-specific datasets and…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Arnav Singhvi , Vasiliki Bikia , Asad Aali , Akshay Chaudhari , Roxana Daneshjou

Safeguarding vision-language models (VLMs) is a critical challenge, as existing methods often suffer from over-defense, which harms utility, or rely on shallow alignment, failing to detect complex threats that require deep reasoning. To…

密码学与安全 · 计算机科学 2026-04-03 Nanxi Li , Zhengyue Zhao , G. Edward Suh , Marco Pavone , Chaowei Xiao

This report presents our team's technical solution for participating in Track 3 of the 2024 ECCV ROAD++ Challenge. The task of Track 3 is atomic activity recognition, which aims to identify 64 types of atomic activities in road scenes based…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Ruyang Li , Tengfei Zhang , Heng Zhang , Tiejun Liu , Yanwei Wang , Xuelei Li

This report provides an architecture-led analysis of two modern vision-language models (VLMs), Qwen2.5-VL-7B-Instruct and Llama-4-Scout-17B-16E-Instruct, and explains how their architectural properties map to a practical video-to-artifact…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Thomson Tong , Diba Darooneh

After decades of progress and effort, obtaining a phase diagram for a strongly-correlated topological system still remains a challenge. Although in principle one could turn to Wilson loops and long-range entanglement, evaluating these…

强关联电子 · 物理学 2017-12-14 Yi Zhang , Roger G. Melko , Eun-Ah Kim

Since the traffic administration at road intersections determines the capacity bottleneck of modern transportation systems, intelligent cooperative coordination for connected autonomous vehicles (CAVs) has shown to be an effective solution.…

机器人学 · 计算机科学 2023-11-14 Donglin Li , Tingting Zhang , Jiping Luo , Tianhao Liang , Bin Cao , Xuanli Wu , Qinyu Zhang

Pre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities. When CLIP is used for depth estimation tasks, the patches,…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Xueting Hu , Ce Zhang , Yi Zhang , Bowen Hai , Ke Yu , Zhihai He

Urban flooding poses an escalating threat to transportation network continuity, yet no operational system currently provides real-time, street-level flood depth information at the centimeter resolution required for dynamic routing, electric…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Nafis Fuad , Xiaodong Qian

Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language Models (VLMs), researchers are exploring their use for tasks such as detecting safety rule…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Xuezheng Chen , Zhengbo Zou

We focus on the problem of detecting traffic events in a surveillance scenario, including the detection of both vehicle actions and traffic collisions. Existing event detection systems are mostly learning-based and have achieved convincing…

计算机视觉与模式识别 · 计算机科学 2020-02-04 Lijun Yu , Peng Chen , Wenhe Liu , Guoliang Kang , Alexander G. Hauptmann