English

Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations

Robotics 2026-07-17 v1

Abstract

Classical leader-follower formation control suffers from single points of failure and error propagation, and relies on absolute localization sensors that are ill-suited for GPS-denied environments. We address these limitations by introducing a fully decentralized, vision-only relative pose estimation framework based on Graph Neural Networks (GNNs). The key idea is the implicit virtual leader (IVL): a non-physical formation reference frame that is not tied to any individual robot but is implicitly learned within the GNN using only monocular images and inter-robot communication. We attach a heteroscedastic GNLL head for aleatoric uncertainty and MC~Dropout for epistemic uncertainty, and conduct a systematic comparison across simulation and real-world test sets. Our framework achieves competitive pose estimation accuracy and generalizes naturally to heterogeneous robot platforms and varying formation sizes.

Cite

@article{arxiv.2607.15708,
  title  = {Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations},
  author = {Shiyuan Yang and Zelin Wang and Zhijia Tao and Yilin Wang and Zhengyu Hou and Xiaosong Kong and Borong Zhang and Yip Fun Yeung and Yuankai Luo and Sharon Lee and Qingbiao Li},
  journal= {arXiv preprint arXiv:2607.15708},
  year   = {2026}
}