中文

HelloMeme:集成空间编织注意力以嵌入高层次且保真信息丰富的扩散模型

计算机视觉与模式识别 2024-10-31 v1

摘要

我们提出了一种有效的方法,将适配器插入文本到图像基础模型中,以实现复杂下游任务的同时保持基础模型的泛化能力。该方法的核心思想是优化与 2D 特征图相关的注意力机制,从而增强适配器的性能。该方法在 meme 视频生成任务上取得了显著成果。我们希望本工作为大型文本到图像模型的后训练任务提供思路。此外,由于该方法与 SD1.5 派生模型兼容性良好,对开源社区具有一定价值。因此,我们将发布相关代码(\url{https://songkey.github.io/hellomeme})。

关键词

引用

@article{arxiv.2410.22901,
  title  = {HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models},
  author = {Shengkai Zhang and Nianhong Jiao and Tian Li and Chaojie Yang and Chenhui Xue and Boya Niu and Jun Gao},
  journal= {arXiv preprint arXiv:2410.22901},
  year   = {2024}
}

备注

11 pages, 7 figures, 2 tables