Clust-PSI-PFL:基于聚类的非 IID 个人化联合学习中的人口稳定性指数方法
摘要
联合学习支持通过保持数据在客户端设备上来实现隐私保护的分布式机器学习模型训练。然而,跨客户端的非独立且不完全相同(non-IID)数据会偏斜更新并降低性能。为了缓解这些问题,我们提出 Clust-PSI-PFL,这是一个基于聚类的个人化联合学习框架,使用人口稳定性指数(PSI)来量化非 IID 数据的程度。我们计算加权 PSI 指标 ,我们证明其比常见的非 IID 指标(Hellinger、Jensen-Shannon 和 Earth Mover's distance)更具信息量。使用 PSI 特征,我们通过 K-means++ 对客户端形成分布均匀的组别,该选定最佳聚类数的标准基于系统的 Silhouette 程序,通常仅产生少数聚类且开销适中。 Across six datasets (tabular, image, and text modalities), two partition protocols (Dirichlet with parameter and Similarity with parameter S), and multiple client sizes, Clust-PSI-PFL delivers up to 18% higher global accuracy than state-of-the-art baselines and markedly improves client fairness by a relative improvement of 37% under severe non-IID data. These results establish PSI-guided clustering as a principled, lightweight mechanism for robust PFL under label skew.
引用
@article{arxiv.2512.20363,
title = {Clust-PSI-PFL: A Population Stability Index Approach for Clustered Non-IID Personalized Federated Learning},
author = {Daniel M. Jimenez-Gutierrez and Mehrdad Hassanzadeh and David Solans and Mohammed Elbamby and Nicolas Kourtellis and Aris Anagnostopoulos and Ioannis Chatzigiannakis and Andrea Vitaletti},
journal= {arXiv preprint arXiv:2512.20363},
year = {2026}
}
备注
Accepted for publication to the 40th IEEE International Parallel & Distributed Processing Symposium (IPDPS 2026)