English
Related papers

Related papers: V2X-UniPool: Unifying Multimodal Perception and Kn…

200 papers

Unified Vision-Language Models (UVLMs) aim to advance multimodal learning by supporting both understanding and generation within a single framework. However, existing approaches largely focus on architectural unification while overlooking…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Shengqiong Wu , Bobo Li , Xinkai Wang , Xiangtai Li , Lei Cui , Furu Wei , Shuicheng Yan , Hao Fei , Tat-seng Chua

Autonomous driving faces safety challenges due to a lack of global perspective and the semantic information of vectorized high-definition (HD) maps. Information from roadside cameras can greatly expand the map perception range through…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Miao Fan , Shanshan Yu , Shengtong Xu , Kun Jiang , Haoyi Xiong , Xiangzeng Liu

Connected and Automated Vehicles use sensors and wireless communication to improve road safety and efficiency. However, attackers may target Vehicle-to-Everything communication. Indeed, an attacker may send authenticated but wrong data to…

Cryptography and Security · Computer Science 2021-12-07 Mohammad Raashid Ansari , Jean-Philippe Monteuuis , Jonathan Petit , Cong Chen

Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation? To investigate this, we introduce UniSandbox, a decoupled evaluation framework…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yuwei Niu , Weiyang Jin , Jiaqi Liao , Chaoran Feng , Peng Jin , Bin Lin , Zongjian Li , Bin Zhu , Weihao Yu , Li Yuan

With the advent of sixth-generation (6G) mobile communication technology, vehicle-to-everything (V2X) communication faces unprecedented challenges in communication efficiency, system generalization capabilities, and model collaboration.…

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-level traffic scene…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Rui Zhao , Qirui Yuan , Jinyu Li , Haofeng Hu , Yun Li , Chengyuan Zheng , Fei Gao

Multimodal large language models (MLLMs) often struggle to ground reasoning in perceptual evidence. We present a systematic study of perception strategies-explicit, implicit, visual, and textual-across four multimodal benchmarks and two…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yizhuo Ding , Mingkang Chen , Zhibang Feng , Tong Xiao , Wanying Qu , Wenqi Shao , Yanwei Fu

Recent advancements in Vehicle-to-Everything communication technology have enabled autonomous vehicles to share sensory information to obtain better perception performance. With the rapid growth of autonomous vehicles and intelligent…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Hao Xiang , Runsheng Xu , Xin Xia , Zhaoliang Zheng , Bolei Zhou , Jiaqi Ma

As self-driving cars increasingly penetrate our daily lives, vehicle-to-everything (V2X) communications are emerging as one of the key enabler technologies. However, the dynamicity of vehicles (one of whose causes is the mobility of…

Systems and Control · Electrical Eng. & Systems 2024-01-30 Dhruba Sunuwar , Seungmo Kim

Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inputs as deterministic, overlooking the perceptual ambiguity…

Artificial Intelligence · Computer Science 2026-01-16 Rui Liu , Dian Yu , Tong Zheng , Runpeng Dai , Zongxia Li , Wenhao Yu , Zhenwen Liang , Linfeng Song , Haitao Mi , Pratap Tokekar , Dong Yu

Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements like lane centerlines and their topology. Existing lane topology reasoning methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Han Li , Yulu Gao , Si Liu , Yuhang Wang , Bo Liu , Beipeng Mu

Simulation is a crucial step in ensuring accurate, efficient, and realistic Connected and Autonomous Vehicles (CAVs) testing and validation. As the adoption of CAV accelerates, the integration of real-world data into simulation environments…

Robotics · Computer Science 2024-09-27 Junwei You , Pei Li , Yang Cheng , Keshu Wu , Rui Gan , Steven T. Parker , Bin Ran

With the development of autonomous driving, it is becoming increasingly common for autonomous vehicles (AVs) and human-driven vehicles (HVs) to travel on the same roads. Existing single-vehicle planning algorithms on board struggle to…

Robotics · Computer Science 2023-02-15 Licheng Wen , Pinlong Cai , Daocheng Fu , Song Mao , Yikang Li

Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything (V2X) communication to share perceptual data, enhances detection but is vulnerable to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Xi Zhou , Tao Huang , Qing-Long Han , Rana Abbas , Mostafa Rahimi Azghadi

The application of computer vision is gradually increasing across various domains. They employ deep learning models with a black-box nature. Without the ability to explain the behavior of neural networks, especially their decision-making…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Maryam Sadat Hosseini Azad , Shahriar Baradaran Shokouhi , Amir Abbas Hamidi Imani , Shahin Atakishiyev , Randy Goebel

Vision-centric autonomous driving has demonstrated excellent performance with economical sensors. As the fundamental step, 3D perception aims to infer 3D information from 2D images based on 3D-2D projection. This makes driving perception…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ye Li , Wenzhao Zheng , Xiaonan Huang , Kurt Keutzer

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven and rule-based, have been hindered by their inability to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Licheng Wen , Xuemeng Yang , Daocheng Fu , Xiaofeng Wang , Pinlong Cai , Xin Li , Tao Ma , Yingxuan Li , Linran Xu , Dengke Shang , Zheng Zhu , Shaoyan Sun , Yeqi Bai , Xinyu Cai , Min Dou , Shuanglu Hu , Botian Shi , Yu Qiao

Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for…

Autonomous driving (AD) systems are becoming increasingly capable of handling complex tasks, mainly due to recent advances in deep learning and AI. As interactions between autonomous systems and humans increase, the interpretability of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Mukilan Karuppasamy , Shankar Gangisetty , Shyam Nandan Rai , Carlo Masone , C V Jawahar

Ensuring safe, comfortable, and efficient planning is crucial for autonomous driving systems. While end-to-end models trained on large datasets perform well in standard driving scenarios, they struggle with complex low-frequency events.…

‹ Prev 1 8 9 10 Next ›