English

S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to Reality

Computer Vision and Pattern Recognition 2024-02-22 v4

Abstract

Due to the lack of enough real multi-agent data and time-consuming of labeling, existing multi-agent cooperative perception algorithms usually select the simulated sensor data for training and validating. However, the perception performance is degraded when these simulation-trained models are deployed to the real world, due to the significant domain gap between the simulated and real data. In this paper, we propose the first Simulation-to-Reality transfer learning framework for multi-agent cooperative perception using a novel Vision Transformer, named as S2R-ViT, which considers both the Deployment Gap and Feature Gap between simulated and real data. We investigate the effects of these two types of domain gaps and propose a novel uncertainty-aware vision transformer to effectively relief the Deployment Gap and an agent-based feature adaptation module with inter-agent and ego-agent discriminators to reduce the Feature Gap. Our intensive experiments on the public multi-agent cooperative perception datasets OPV2V and V2V4Real demonstrate that the proposed S2R-ViT can effectively bridge the gap from simulation to reality and outperform other methods significantly for point cloud-based 3D object detection.

Keywords

Cite

@article{arxiv.2307.07935,
  title  = {S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to Reality},
  author = {Jinlong Li and Runsheng Xu and Xinyu Liu and Baolu Li and Qin Zou and Jiaqi Ma and Hongkai Yu},
  journal= {arXiv preprint arXiv:2307.07935},
  year   = {2024}
}

Comments

submit the latest one, accepted by the 2024 IEEE International Conference on Robotics and Automation (ICRA)

R2 v1 2026-06-28T11:31:32.575Z