中文
相关论文

相关论文: ChatBEV: A Visual Language Model that Understands …

200 篇论文

Generating a detailed near-field perceptual model of the environment is an important and challenging problem in both self-driving vehicles and autonomous mobile robotics. A Bird Eye View (BEV) map, providing a panoptic representation, is a…

计算机视觉与模式识别 · 计算机科学 2022-06-01 Pramit Dutta , Ganesh Sistu , Senthil Yogamani , Edgar Galván , John McDonald

Autonomous Vehicles (AVs) are transforming the future of transportation through advances in intelligent perception, decision-making, and control systems. However, their success is tied to one core capability, reliable object detection in…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Sayed Pedram Haeri Boroujeni , Niloufar Mehrabi , Hazim Alzorgan , Mahlagha Fazeli , Abolfazl Razi

Online scene perception and topology reasoning are critical for autonomous vehicles to understand their driving environments, particularly for mapless driving systems that endeavor to reduce reliance on costly High-Definition (HD) maps.…

机器人学 · 计算机科学 2025-06-27 Muleilan Pei , Jiayao Shan , Peiliang Li , Jieqi Shi , Jing Huo , Yang Gao , Shaojie Shen

Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Zhaowei Wang , Hongming Zhang , Tianqing Fang , Ye Tian , Yue Yang , Kaixin Ma , Xiaoman Pan , Yangqiu Song , Dong Yu

Bird's eye view (BEV) semantic segmentation plays a crucial role in spatial sensing for autonomous driving. Although recent literature has made significant progress on BEV map understanding, they are all based on single-agent camera-based…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Runsheng Xu , Zhengzhong Tu , Hao Xiang , Wei Shao , Bolei Zhou , Jiaqi Ma

Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happening within the scene for safe interaction with other road…

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets…

Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi-sensor…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Mingzhe Tao , Ruiping Liu , Junwei Zheng , Yufan Chen , Kedi Ying , M. Saquib Sarfraz , Kailun Yang , Jiaming Zhang , Rainer Stiefelhagen

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Xin Zhou , Dingkang Liang , Sifan Tu , Xiwu Chen , Yikang Ding , Dingyuan Zhang , Feiyang Tan , Hengshuang Zhao , Xiang Bai

Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visual evidence drives these judgments. We study whether…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Carlos Hinojosa , Clemens Grange , Bernard Ghanem

In autonomous driving, high-definition (HD) maps and semantic maps in bird's-eye view (BEV) are essential for accurate localization, planning, and decision-making. This paper introduces an enhanced End-to-End model named MapFM for online…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Leonid Ivanov , Vasily Yuryev , Dmitry Yudin

Understanding the road genome is essential to realize autonomous driving. This highly intelligent problem contains two aspects - the connection relationship of lanes, and the assignment relationship between lanes and traffic elements, where…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Tianyu Li , Li Chen , Huijie Wang , Yang Li , Jiazhi Yang , Xiangwei Geng , Shengyin Jiang , Yuting Wang , Hang Xu , Chunjing Xu , Junchi Yan , Ping Luo , Hongyang Li

Autonomous driving faces critical challenges in rare long-tail events and complex multi-agent interactions, which are scarce in real-world data yet essential for robust safety validation. This paper presents a high-fidelity scenario…

机器学习 · 计算机科学 2025-11-27 Yuhang Wang , Heye Huang , Zhenhua Xu , Kailai Sun , Baoshen Guo , Jinhua Zhao

To assist human drivers and autonomous vehicles in assessing crash risks, driving scene analysis using dash cameras on vehicles and deep learning algorithms is of paramount importance. Although these technologies are increasingly available,…

计算机视觉与模式识别 · 计算机科学 2021-06-22 Muhammad Monjurul Karim , Yu Li , Ruwen Qin , Zhaozheng Yin

ChatGPT embarks on a new era of artificial intelligence and will revolutionize the way we approach intelligent traffic safety systems. This paper begins with a brief introduction about the development of large language models (LLMs). Next,…

计算与语言 · 计算机科学 2023-09-07 Ou Zheng , Mohamed Abdel-Aty , Dongdong Wang , Zijin Wang , Shengxuan Ding

The integration of Large Language Models (LLMs) with computer vision is profoundly transforming perception tasks like image segmentation. For intelligent transportation systems (ITS), where accurate scene understanding is critical for…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Sanjeda Akter , Ibne Farabi Shihab , Anuj Sharma

Connected autonomous vehicles (CAVs) must simultaneously perform multiple tasks, such as object detection, semantic segmentation, depth estimation, trajectory prediction, motion prediction, and behaviour prediction, to ensure safe and…

机器人学 · 计算机科学 2025-08-07 Jiayuan Wang , Farhad Pourpanah , Q. M. Jonathan Wu , Ning Zhang

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet most evaluations of…

We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object…

Lane segment topology reasoning provides comprehensive bird's-eye view (BEV) road scene understanding, which can serve as a key perception module in planning-oriented end-to-end autonomous driving systems. Existing lane topology reasoning…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Yiming Yang , Hongbin Lin , Yueru Luo , Suzhong Fu , Chao Zheng , Xinrui Yan , Shuqi Mei , Kun Tang , Shuguang Cui , Zhen Li