English

Intrinsic-Motivation Multi-Robot Social Formation Navigation with Coordinated Exploration

Robotics 2025-12-17 v2 Artificial Intelligence

Abstract

This paper investigates the application of reinforcement learning (RL) to multi-robot social formation navigation, a critical capability for enabling seamless human-robot coexistence. While RL offers a promising paradigm, the inherent unpredictability and often uncooperative dynamics of pedestrian behavior pose substantial challenges, particularly concerning the efficiency of coordinated exploration among robots. To address this, we propose a novel coordinated-exploration multi-robot RL algorithm introducing an intrinsic motivation exploration. Its core component is a self-learning intrinsic reward mechanism designed to collectively alleviate policy conservatism. Moreover, this algorithm incorporates a dual-sampling mode within the centralized training and decentralized execution framework to enhance the representation of both the navigation policy and the intrinsic reward, leveraging a two-time-scale update rule to decouple parameter updates. Empirical results on social formation navigation benchmarks demonstrate the proposed algorithm's superior performance over existing state-of-the-art methods across crucial metrics. Our code and video demos are available at: https://github.com/czxhunzi/CEMRRL.

Keywords

Cite

@article{arxiv.2512.13293,
  title  = {Intrinsic-Motivation Multi-Robot Social Formation Navigation with Coordinated Exploration},
  author = {Hao Fu and Wei Liu and Shuai Zhou},
  journal= {arXiv preprint arXiv:2512.13293},
  year   = {2025}
}
R2 v1 2026-07-01T08:25:12.529Z