English

Distilling Cross-Modal Knowledge via Feature Disentanglement

Computer Vision and Pattern Recognition 2025-11-26 v1 Artificial Intelligence

Abstract

Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead to difficult knowledge transfer. To address this challenge, we propose frequency-decoupled cross-modal knowledge distillation, a method designed to decouple and balance knowledge transfer across modalities by leveraging frequency-domain features. We observed that low-frequency features exhibit high consistency across different modalities, whereas high-frequency features demonstrate extremely low cross-modal similarity. Accordingly, we apply distinct losses to these features: enforcing strong alignment in the low-frequency domain and introducing relaxed alignment for high-frequency features. We also propose a scale consistency loss to address distributional shifts between modalities, and employ a shared classifier to unify feature spaces. Extensive experiments across multiple benchmark datasets show our method substantially outperforms traditional KD and state-of-the-art cross-modal KD approaches. Code is available at https://github.com/Johumliu/FD-CMKD.

Keywords

Cite

@article{arxiv.2511.19887,
  title  = {Distilling Cross-Modal Knowledge via Feature Disentanglement},
  author = {Junhong Liu and Yuan Zhang and Tao Huang and Wenchao Xu and Renyu Yang},
  journal= {arXiv preprint arXiv:2511.19887},
  year   = {2025}
}

Comments

Accepted by AAAI 2026

R2 v1 2026-07-01T07:53:30.932Z