中文
相关论文

相关论文: Multi-Stage VLM Pipeline for Zero-Shot Traffic Acc…

200 篇论文

Existing view planning systems either adopt an iterative paradigm using next-best views (NBV) or a one-shot pipeline relying on the set-covering view-planning (SCVP) network. However, neither of these methods can concurrently guarantee both…

机器人学 · 计算机科学 2024-10-31 Sicong Pan , Hao Hu , Hui Wei , Nils Dengler , Tobias Zaenker , Murad Dawood , Maren Bennewitz

Prompt ensembling of Large Language Model (LLM) generated category-specific prompts has emerged as an effective method to enhance zero-shot recognition ability of Vision-Language Models (VLMs). To obtain these category-specific prompts, the…

计算机视觉与模式识别 · 计算机科学 2024-08-08 M. Jehanzeb Mirza , Leonid Karlinsky , Wei Lin , Sivan Doveh , Jakub Micorek , Mateusz Kozinski , Hilde Kuehne , Horst Possegger

The claim matching (CM) task can benefit an automated fact-checking pipeline by putting together claims that can be resolved with the same fact-check. In this work, we are the first to explore zero-shot and few-shot learning approaches to…

计算与语言 · 计算机科学 2025-03-04 Dina Pisarevskaya , Arkaitz Zubiaga

Bridges, as critical components of civil infrastructure, are increasingly affected by deterioration, making reliable traffic monitoring essential for assessing their remaining service life. Among operational loads, traffic load plays a…

机器学习 · 计算机科学 2026-01-21 Hanshuo Wu , Xudong Jian , Christos Lataniotis , Cyprien Hoelzl , Eleni Chatzi , Yves Reuland

Virtual Traffic Light (VTL) is a traffic control method that does not require traffic signal-related infrastructure for roadway intersections. Connected vehicles (CVs) are given right-of-way based on prevailing traffic conditions, such as…

其他计算机科学 · 计算机科学 2025-06-06 Abyad Enan , M Sabbir Salek , Mashrur Chowdhury , Gurcan Comert , Sakib M. Khan , Reek Majumder

For the majority of the machine learning community, the expensive nature of collecting high-quality human-annotated data and the inability to efficiently finetune very large state-of-the-art pretrained models on limited compute are major…

计算机视觉与模式识别 · 计算机科学 2022-11-07 Anuj Diwan , Puyuan Peng , Raymond J. Mooney

Employing Vehicle-to-Vehicle communication to enhance perception performance in self-driving technology has attracted considerable attention recently; however, the absence of a suitable open dataset for benchmarking algorithms has made it…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Runsheng Xu , Hao Xiang , Xin Xia , Xu Han , Jinlong Li , Jiaqi Ma

This thesis considers the problem of scheduling autonomous vehicles at intersections. A new system is proposed which is more efficient and could replace the recently introduced Autonomous Intersection Management (AIM) model. The proposed…

其他计算机科学 · 计算机科学 2018-05-17 Nasser Aloufi

This technical report presents the 2nd winning model for AQTC, a task newly introduced in CVPR 2022 LOng-form VidEo Understanding (LOVEU) challenges. This challenge faces difficulties with multi-step answers, multi-modal, and diverse and…

计算机视觉与模式识别 · 计算机科学 2022-06-30 Hyeonyu Kim , Jongeun Kim , Jeonghun Kang , Sanguk Park , Dongchan Park , Taehwan Kim

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigorous biometric tasks remains unexplored. This work presents an exploratory study…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Marta Robledo-Moreno , Ruben Vera-Rodriguez , Ruben Tolosana , Javier Ortega-Garcia

Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these tasks are handled separately, with linguistic interpretation left to human experts. This paper…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Malikussaid , Imad Gohar

Landing safety is a challenge heavily engaging the research community recently, due to the increasing interest in applications availed by aerial vehicles. In this paper, we propose a landing safety pipeline based on state of the art object…

机器人学 · 计算机科学 2024-04-09 Tilemahos Mitroudas , Vasiliki Balaska , Athanasios Psomoulis , Antonios Gasteratos

Purpose: Accurate detection and 6D pose estimation of surgical instruments are crucial for many computer-assisted interventions. However, supervised methods lack flexibility for new or unseen tools and require extensive annotated data. This…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Jonas Hein , Lilian Calvet , Matthias Seibold , Siyu Tang , Marc Pollefeys , Philipp Fürnstahl

Large-scale pre-trained models have demonstrated impressive performance in vision and language tasks within open-world scenarios. Due to the lack of comparable pre-trained models for 3D shapes, recent methods utilize language-image…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Dan Song , Xinwei Fu , Ning Liu , Weizhi Nie , Wenhui Li , Lanjun Wang , You Yang , Anan Liu

Safety hazard identification and prevention are the key elements of proactive safety management. Previous research has extensively explored the applications of computer vision to automatically identify hazards from image clips collected…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Muhammad Adil , Gaang Lee , Vicente A. Gonzalez , Qipei Mei

Automatic license plate recognition (ALPR) and vehicle make and model recognition underpin intelligent transportation systems, supporting law enforcement, toll collection, and post-incident investigation. Applying these methods to videos…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Pouya Parsa , Keya Li , Kara M. Kockelman , Seongjin Choi

Accurate 6-DoF object pose estimation and tracking are critical for reliable robotic manipulation. However, zero-shot methods often fail under viewpoint-induced ambiguities and fixed-camera setups struggle when objects move or become…

机器人学 · 计算机科学 2026-03-10 Sheng Liu , Zhe Li , Weiheng Wang , Han Sun , Heng Zhang , Hongpeng Chen , Yusen Qin , Arash Ajoudani , Yizhao Wang

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint…

机器人学 · 计算机科学 2025-06-18 Xiangyu Shi , Zerui Li , Wenqi Lyu , Jiatong Xia , Feras Dayoub , Yanyuan Qiao , Qi Wu

Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, is a critical challenge for autonomous driving systems, as crashes involving VRUs often result in severe or fatal consequences. While multimodal large…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Younggun Kim , Ahmed S. Abdelrahman , Mohamed Abdel-Aty

Accurate point-wise velocity estimation in 3D is crucial for robot interaction with non-rigid dynamic agents, enabling robust performance in path planning, collision avoidance, and object manipulation in dynamic environments. To this end,…

机器人学 · 计算机科学 2026-04-13 Landson Guo , Andres M. Diaz Aguilar , William Talbot , Turcan Tuna , Marco Hutter , Cesar Cadena