中文
相关论文

相关论文: CoST: Efficient Collaborative Perception From Unif…

200 篇论文

Cooperative perception is a promising technique for intelligent and connected vehicles through vehicle-to-everything (V2X) cooperation, provided that accurate pose information and relative pose transforms are available. Nevertheless,…

机器人学 · 计算机科学 2024-02-23 Zhiying Song , Tenghui Xie , Hailiang Zhang , Jiaxin Liu , Fuxi Wen , Jun Li

Collaborative perception in Internet of Vehicles (IoV) aggregates multi-vehicle observations for broader scene coverage and improved decision-making. However, fusion quality degrades under spatiotemporal heterogeneity from unsynchronized…

网络与互联网体系结构 · 计算机科学 2026-02-17 Qiaomei Han , Xianbin Wang , Minghui Liwang , Dusit Niyato

Collaborative perception improves 3D understanding by fusing multi-agent observations, yet intermediate-feature sharing faces strict bandwidth constraints as dense BEV features saturate V2X links. We observe that collaborators view the same…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yuankun Zeng , Shaohui Li , Zhi Li , Shulan Ruan , Yu Liu , You He

Comprehensive environment perception is essential for autonomous vehicles to operate safely. It is crucial to detect both dynamic road users and static objects like traffic signs or lanes as these are required for safe motion planning.…

机器人学 · 计算机科学 2025-12-17 Jörg Gamerdinger , Sven Teufel , Georg Volk , Oliver Bringmann

Occlusion is a major challenge for LiDAR-based object detection methods. This challenge becomes safety-critical in urban traffic where the ego vehicle must have reliable object detection to avoid collision while its field of view is…

机器人学 · 计算机科学 2023-09-20 Minh-Quan Dao , Julie Stephany Berrio , Vincent Frémont , Mao Shan , Elwan Héry , Stewart Worrall

Camera and radar sensors have significant advantages in cost, reliability, and maintenance compared to LiDAR. Existing fusion methods often fuse the outputs of single modalities at the result-level, called the late fusion strategy. This can…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Youngseok Kim , Sanmin Kim , Jun Won Choi , Dongsuk Kum

Devising intelligent agents able to live in an environment and learn by observing the surroundings is a longstanding goal of Artificial Intelligence. From a bare Machine Learning perspective, challenges arise when the agent is prevented…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Matteo Tiezzi , Simone Marullo , Lapo Faggi , Enrico Meloni , Alessandro Betti , Stefano Melacci

Vehicle-to-vehicle (V2V) communications have greatly enhanced the perception capabilities of connected and automated vehicles (CAVs) by enabling information sharing to "see through the occlusions", resulting in significant performance…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Yunsheng Ma , Juanwu Lu , Can Cui , Sicheng Zhao , Xu Cao , Wenqian Ye , Ziran Wang

Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture,…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Dongho Lee , Jongseo Lee , Jinwoo Choi

Multi-modality is an important feature of sensor based activity recognition. In this work, we consider two inherent characteristics of human activities, the spatially-temporally varying salience of features and the relations between…

人机交互 · 计算机科学 2019-05-23 Kaixuan Chen , Lina Yao , Dalin Zhang , Bin Guo , Zhiwen Yu

Collaborative Perception (CP) is a process in which an ego agent receives and fuses sensor information from surrounding vehicles and infrastructure to enhance its perception capability. To evaluate the need for infrastructure equipped with…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Hyunchul Bae , Minhee Kang , Minwoo Song , Heejin Ahn

Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most of the existing methods only model the temporal information and the spatial information for videos in an independent…

计算机视觉与模式识别 · 计算机科学 2022-04-21 Zheng Chang , Xinfeng Zhang , Shanshe Wang , Siwei Ma , Wen Gao

The major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that attention mechanism is beneficial to the fusion process. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2020-08-18 Bin Duan , Hao Tang , Wei Wang , Ziliang Zong , Guowei Yang , Yan Yan

While significant advances have been made for single-agent perception, many applications require multiple sensing agents and cross-agent communication due to benefits such as coverage and robustness. It is therefore critical to develop…

计算机视觉与模式识别 · 计算机科学 2020-06-04 Yen-Cheng Liu , Junjiao Tian , Nathaniel Glaser , Zsolt Kira

Comprehensive perception of the environment is crucial for the safe operation of autonomous vehicles. However, the perception capabilities of autonomous vehicles are limited due to occlusions, limited sensor ranges, or environmental…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Sven Teufel , Jörg Gamerdinger , Georg Volk , Oliver Bringmann

In the past decade, although single-robot perception has made significant advancements, the exploration of multi-robot collaborative perception remains largely unexplored. This involves fusing compressed, intermittent, limited,…

机器人学 · 计算机科学 2024-05-24 Yang Zhou , Long Quang , Carlos Nieto-Granda , Giuseppe Loianno

Vision and touch are two of the important sensing modalities for humans and they offer complementary information for sensing the environment. Robots could also benefit from such multi-modal sensing ability. In this paper, addressing for the…

机器人学 · 计算机科学 2018-03-14 Shan Luo , Wenzhen Yuan , Edward Adelson , Anthony G. Cohn , Raul Fuentes

In the era of information explosion, spatio-temporal data mining serves as a critical part of urban management. Considering the various fields demanding attention, e.g., traffic state, human activity, and social event, predicting multiple…

人工智能 · 计算机科学 2023-09-19 Zijian Zhang , Xiangyu Zhao , Qidong Liu , Chunxu Zhang , Qian Ma , Wanyu Wang , Hongwei Zhao , Yiqi Wang , Zitao Liu

In this paper we introduce Co-Fusion, a dense SLAM system that takes a live stream of RGB-D images as input and segments the scene into different objects (using either motion or semantic cues) while simultaneously tracking and…

计算机视觉与模式识别 · 计算机科学 2017-09-06 Martin Rünz , Lourdes Agapito

Generative models are widely utilized to model the distribution of fused images in the field of infrared and visible image fusion. However, current generative models based fusion methods often suffer from unstable training and slow…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Zhiming Meng , Hui Li , Zeyang Zhang , Zhongwei Shen , Yunlong Yu , Xiaoning Song , Xiaojun Wu