中文
相关论文

相关论文: V2X-UniPool: Unifying Multimodal Perception and Kn…

200 篇论文

Vehicle-to-everything technologies (V2X) have become an ideal paradigm to extend the perception range and see through the occlusion. Exiting efforts focus on single-frame cooperative perception, however, how to capture the temporal cue…

机器学习 · 计算机科学 2025-11-04 Xinyu Zhang , Zewei Zhou , Zhaoyi Wang , Yangjie Ji , Yanjun Huang , Hong Chen

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers…

Vehicle-to-Everything (V2X) collaborative perception is crucial for autonomous driving. However, achieving high-precision V2X perception requires a significant amount of annotated real-world data, which can always be expensive and hard to…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Xianghao Kong , Wentao Jiang , Jinrang Jia , Yifeng Shi , Runsheng Xu , Si Liu

Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations,…

Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-thought interleaved with visual latent tokens. This interleaved…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Houcheng Jiang , Jiajun Fu , Junfeng Fang , Chen Gao , Xiang Wang , Xiangnan He , Yong Li

Vision Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning could enhance temporal consistency and perception action…

Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zimo Wen , Boxiu Li , Wanbo Zhang , Junxiang Lei , Xiaoyu Chen , Yijia Fan , Qi Zhang , Yujiang Wang , Lili Qiu , Bo Li , Ziwei Liu , Caihua Shan , Yifan Yang , Yifei Shen

The ability to perceive and comprehend a traffic situation and to estimate the state of the vehicles and road-users in the surrounding of the ego-vehicle is known as situational awareness. Situational awareness for a heavy-duty autonomous…

机器人学 · 计算机科学 2023-05-30 Vandana Narri , Amr Alanwar , Jonas Mårtensson , Christoffer Norén , Karl Henrik Johansson

Modern autonomous vehicle perception systems are often constrained by occlusions, blind spots, and limited sensing range. While existing cooperative perception paradigms, such as Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure (V2I),…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Weijia Li , Haoen Xiang , Tianxu Wang , Shuaibing Wu , Qiming Xia , Cheng Wang , Chenglu Wen

The development of high-level autonomous driving (AD) is shifting from perception-centric limitations to a more fundamental bottleneck, namely, a deficit in robust and generalizable reasoning. Although current AD systems manage structured…

人工智能 · 计算机科学 2026-03-13 Kejin Yu , Yuhan Sun , Taiqiang Wu , Ruixu Zhang , Zhiqiang Lin , Yuxin Meng , Junjie Wang , Yujiu Yang

Vehicle-to-everything (V2X) perception refers to a suite of technologies that empower vehicles to sense their environment and communicate with other entities, including surrounding vehicles, infrastructure, and cloud/edge networks. With the…

Vision-language models (VLMs) have emerged as a promising direction for end-to-end autonomous driving (AD) by jointly modeling visual observations, driving context, and language-based reasoning. However, existing VLM-based systems face a…

机器人学 · 计算机科学 2026-03-10 Ximeng Tao , Pardis Taghavi , Dimitar Filev , Reza Langari , Gaurav Pandey

To ensure safe operation of autonomous vehicles in complex urban environments, complete perception of the environment is necessary. However, due to environmental conditions, sensor limitations, and occlusions, this is not always possible…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Sven Teufel , Jörg Gamerdinger , Jan-Patrick Kirchner , Georg Volk , Oliver Bringmann

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Xinyang Wang , Qian Liu , Wenjie Ding , Zhao Yang , Wei Li , Chang Liu , Bailin Li , Kun Zhan , Xianpeng Lang , Wei Chen

Smart autonomous agents are becoming increasingly important in various real-life applications, including robotics and autonomous vehicles. One crucial skill that these agents must possess is the ability to interact with their surrounding…

机器人学 · 计算机科学 2024-08-09 Niyati Rawal , Roberto Bigazzi , Lorenzo Baraldi , Rita Cucchiara

Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However, current VLM-based approaches suffer from two major…

机器人学 · 计算机科学 2026-03-31 Yuqi Ye , Zijian Zhang , Junhong Lin , Shangkun Sun , Changhao Peng , Wei Gao

The evolution of wireless communications into 6G and beyond is expected to rely on new machine learning (ML)-based capabilities. These can enable proactive decisions and actions from wireless-network components to sustain quality-of-service…

The objective of the collaborative vehicle-to-everything perception task is to enhance the individual vehicle's perception capability through message communication among neighboring traffic agents. Previous methods focus on achieving…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Si Liu , Zihan Ding , Jiahui Fu , Hongyu Li , Siheng Chen , Shifeng Zhang , Xu Zhou

The paper addresses the vehicle-to-X (V2X) data fusion for cooperative or collective perception (CP). This emerging and promising intelligent transportation systems (ITS) technology has enormous potential for improving efficiency and safety…

Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of essential world knowledge to understand and reason about…

机器人学 · 计算机科学 2025-11-05 Yu Gao , Anqing Jiang , Yiru Wang , Wang Jijun , Hao Jiang , Zhigang Sun , Heng Yuwen , Wang Shuo , Hao Zhao , Sun Hao