English
Related papers

Related papers: RefAV: Towards Planning-Centric Scenario Mining

200 papers

Semantic anomalies-context-dependent hazards that pixel-level detectors cannot reason about-pose a critical safety risk in autonomous driving. We propose a \emph{semantic observer layer}: a quantized vision-language model (VLM) running at…

Robotics · Computer Science 2026-04-01 Kunal Runwal , Swaraj Gajare , Daniel Adejumo , Omkar Ankalkope , Siddhant Baroth , Aliasghar Arab

Integrating large language models (LLMs) into embodied AI models is becoming increasingly prevalent. However, existing zero-shot LLM-based Vision-and-Language Navigation (VLN) agents either encode images as textual scene descriptions,…

Artificial Intelligence · Computer Science 2025-09-30 Yue Zhang , Tianyi Ma , Zun Wang , Yanyuan Qiao , Parisa Kordjamshidi

Training and evaluating autonomous driving algorithms requires a diverse range of scenarios. However, most available datasets predominantly consist of normal driving behaviors demonstrated by human drivers, resulting in a limited number of…

Robotics · Computer Science 2025-05-27 Miao Li , Wenhao Ding , Haohong Lin , Yiqi Lyu , Yihang Yao , Yuyou Zhang , Ding Zhao

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Junxiao Xue , Quan Deng , Fei Yu , Yanhao Wang , Jun Wang , Yuehua Li

Ensuring the functional safety of motion planning modules in autonomous vehicles remains a critical challenge, especially when dealing with complex or learning-based software. Online verification has emerged as a promising approach to…

Robotics · Computer Science 2025-07-11 Korbinian Moller , Rafael Neher , Marvin Seegert , Johannes Betz

Autonomous Driving (AD) systems have made notable progress, but their performance in long-tail, safety-critical scenarios remains limited. These rare cases contribute a disproportionate number of accidents. Vision-Language Action (VLA)…

Robotics · Computer Science 2025-09-22 Shiyu Fang , Yiming Cui , Haoyang Liang , Chen Lv , Peng Hang , Jian Sun

This paper presents an integrated motion planning system for autonomous vehicle (AV) parking in the presence of other moving vehicles. The proposed system includes 1) a hybrid environment predictor that predicts the motions of the…

Robotics · Computer Science 2022-04-28 Jessica Leu , Yebin Wang , Masayoshi Tomizuka , Stefano Di Cairano

Safety verification for autonomous vehicles (AVs) and ground robots is crucial for ensuring reliable operation given their uncertain environments. Formal language tools provide a robust and sound method to verify safety rules for such…

Robotics · Computer Science 2025-01-24 Aditya Parameshwaran , Yue Wang

Talk2BEV is a large vision-language model (LVLM) interface for bird's-eye view (BEV) maps in autonomous driving contexts. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set…

Localization is one of the most crucial tasks for Unmanned Aerial Vehicle systems (UAVs) directly impacting overall performance, which can be achieved with various sensors and applied to numerous tasks related to search and rescue…

Robotics · Computer Science 2024-11-05 Thanh Nguyen Canh , Huy-Hoang Ngo , Xiem HoangVan , Nak Young Chong

In the field of Vision-Language Navigation (VLN), aerial datasets remain limited in their ability to combine scale, diversity, and realism, often relying on either costly real-world scenes or visually limited simulations. To address these…

Robotics · Computer Science 2026-05-20 Jinhan Li , Xijie Huang , Zhaoqi Wang , Yijin Wang , Weiqi Ge , Qiyi He , Mo Zhu , Fei Gao , Yuze Wu , Xin Zhou

Multimodal large language models (MLLMs) have shown satisfactory effects in many autonomous driving tasks. In this paper, MLLMs are utilized to solve joint semantic scene understanding and risk localization tasks, while only relying on…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Jiaqi Fan , Jianhua Wu , Jincheng Gao , Jianhao Yu , Yafei Wang , Hongqing Chu , Bingzhao Gao

Automated vehicles (AVs) are expected to increase traffic safety and traffic efficiency, among others by enabling flexible mobility-on-demand systems. This is particularly important in Singapore, being one of the world's most densely…

Robotics · Computer Science 2021-12-20 J. Ploeg , E. de Gelder , M. Slavík , E. Querner , T. Webster , N. de Boer

Safe UAV emergency landing requires more than just identifying flat terrain; it demands understanding complex semantic risks (e.g., crowds, temporary structures) invisible to traditional geometric sensors. In this paper, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Chunliang Hua , Zeyuan Yang , Lei Zhang , Jiayang Sun , Fengwen Chen , Chunlan Zeng , Xiao Hu

The real world is messy and unstructured. Uncovering critical information often requires active, goal-driven exploration. It remains to be seen whether Vision-Language Models (VLMs), which recently emerged as a popular zero-shot tool in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Adam Pardyl , Dominik Matuszek , Mateusz Przebieracz , Marek Cygan , Bartosz Zieliński , Maciej Wołczyk

Autonomous driving technology has the potential to transform transportation, but its wide adoption depends on the development of interpretable and transparent decision-making systems. Scene captioning, which generates natural language…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Felix Brandstaetter , Erik Schuetz , Katharina Winter , Fabian Flohr

With the rise of multimodal applications, instruction data has become critical for training multimodal language models capable of understanding complex image-based queries. Existing practices rely on powerful but costly large language…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Jieyu Zhang , Le Xue , Linxin Song , Jun Wang , Weikai Huang , Manli Shu , An Yan , Zixian Ma , Juan Carlos Niebles , Silvio Savarese , Caiming Xiong , Zeyuan Chen , Ranjay Krishna , Ran Xu

Existing Video Anomaly Detection (VAD) methods typically rely on task-specific training, leading to strong domain dependency and high training costs. Moreover, most existing methods output only scalar anomaly scores, providing limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Hyeongmuk Lim , Youngbum Hur

The development of large vision-language models (LVLMs) offers the potential to address challenges faced by traditional multimodal recommendations thanks to their proficient understanding of static images and textual dynamics. However, the…

Artificial Intelligence · Computer Science 2024-02-14 Yuqing Liu , Yu Wang , Lichao Sun , Philip S. Yu

Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visual evidence drives these judgments. We study whether…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Carlos Hinojosa , Clemens Grange , Bernard Ghanem
‹ Prev 1 8 9 10 Next ›