Vertical Federated Learning (VFL) offers a privacy-preserving paradigm for Edge AI scenarios like mobile health diagnostics, where sensitive multimodal data reside on distributed, resource-constrained devices. Yet, standard VFL systems often suffer performance limitations due to simplistic feature fusion. This paper introduces HybridVFL, a novel framework designed to overcome this bottleneck by employing client-side feature disentanglement paired with a server-side cross-modal transformer for context-aware fusion. Through systematic evaluation on the multimodal HAM10000 skin lesion dataset, we demonstrate that HybridVFL significantly outperforms standard federated baselines, validating the criticality of advanced fusion mechanisms in robust, privacy-preserving systems.
@article{arxiv.2512.10701,
title = {HybridVFL: Disentangled Feature Learning for Edge-Enabled Vertical Federated Multimodal Classification},
author = {Mostafa Anoosha and Zeinab Dehghani and Kuniko Paxton and Koorosh Aslansefat and Dhavalkumar Thakker},
journal= {arXiv preprint arXiv:2512.10701},
year = {2025}
}
Comments
6 pages, 2 figures, 1 table. Accepted at UCC '25 (IEEE/ACM 18th International Conference on Utility and Cloud Computing), December 1-4, 2025, Nantes, France. DOI to be activated upon final publication