English

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

Computer Vision and Pattern Recognition 2026-05-21 v3

Abstract

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this setting, we observe a prompt forgetting phenomenon: the semantics of the prompt representation in the text branch is progressively forgotten as depth increases. We further verify this effect on three representative MMDiTs--SD3, SD3.5, and FLUX.1 by probing linguistic attributes of the representations over the layers in the text branch. Motivated by these findings, we introduce a training-free approach, prompt reinjection, which reinjects prompt representations from early layers into later layers to alleviate this forgetting. Experiments on GenEval, DPG, and T2I-CompBench++ show consistent gains in instruction-following capability, along with improvements on metrics capturing preference, aesthetics, and overall text--image generation quality.

Keywords

Cite

@article{arxiv.2602.06886,
  title  = {Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers},
  author = {Yuxuan Yao and Yuxuan Chen and Hui Li and Kaihui Cheng and Qipeng Guo and Yuwei Sun and Zilong Dong and Jingdong Wang and Siyu Zhu},
  journal= {arXiv preprint arXiv:2602.06886},
  year   = {2026}
}

Comments

19 pages

R2 v1 2026-07-01T10:24:46.866Z