English

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

Computer Vision and Pattern Recognition 2025-12-03 v1

Abstract

Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. While methods using a fixed number of learnable query tokens offer computational efficiency, they suffer from task generalization collapse, failing to adapt to new tasks that are distant from their pre-training tasks. To overcome this, we propose Noisy Query Tokens, which learn a distributed representation space between the VLM and Diffusion Model via end-to-end optimization, enhancing continual learning. Additionally, we introduce a VAE branch with linear projection to recover fine-grained image details. Experimental results confirm our approach mitigates generalization collapse and enables stable continual learning across diverse tasks.

Keywords

Cite

@article{arxiv.2512.02536,
  title  = {WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens},
  author = {Jian Yang and Dacheng Yin and Xiaoxuan He and Yong Li and Fengyun Rao and Jing Lyu and Wei Zhai and Yang Cao and Zheng-Jun Zha},
  journal= {arXiv preprint arXiv:2512.02536},
  year   = {2025}
}