English

FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning

Machine Learning 2025-10-15 v1

Abstract

Text-to-Image (T2I) models have demonstrated their versatility in a wide range of applications. However, adaptation of T2I models to specialized tasks is often limited by the availability of task-specific data due to privacy concerns. On the other hand, harnessing the power of rich multimodal data from modern mobile systems and IoT infrastructures presents a great opportunity. This paper introduces Federated Multi-modal Knowledge Transfer (FedMMKT), a novel framework that enables co-enhancement of a server T2I model and client task-specific models using decentralized multimodal data without compromising data privacy.

Keywords

Cite

@article{arxiv.2510.12254,
  title  = {FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning},
  author = {Ningxin He and Yang Liu and Wei Sun and Xiaozhou Ye and Ye Ouyang and Tiegang Gao and Zehui Zhang},
  journal= {arXiv preprint arXiv:2510.12254},
  year   = {2025}
}
R2 v1 2026-07-01T06:35:51.536Z