English

Measuring Heterogeneity in Machine Learning with Distributed Energy Distance

Machine Learning 2025-01-28 v1 Artificial Intelligence Distributed, Parallel, and Cluster Computing Machine Learning

Abstract

In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensitive measure for quantifying distributional discrepancies. While we show that energy distance is robust for detecting data distribution shifts, its direct use in large-scale systems can be prohibitively expensive. To address this, we develop Taylor approximations that preserve key theoretical quantitative properties while reducing computational overhead. Through simulation studies, we show how accurately capturing feature discrepancies boosts convergence in distributed learning. Finally, we propose a novel application of energy distance to assign penalty weights for aligning predictions across heterogeneous nodes, ultimately enhancing coordination in federated and distributed settings.

Keywords

Cite

@article{arxiv.2501.16174,
  title  = {Measuring Heterogeneity in Machine Learning with Distributed Energy Distance},
  author = {Mengchen Fan and Baocheng Geng and Roman Shterenberg and Joseph A. Casey and Zhong Chen and Keren Li},
  journal= {arXiv preprint arXiv:2501.16174},
  year   = {2025}
}

Comments

15 pages, 5 figures

R2 v1 2026-06-28T21:19:56.158Z