HelloMeme:集成空间编织注意力以嵌入高层次且保真信息丰富的扩散模型
计算机视觉与模式识别
2024-10-31 v1
摘要
我们提出了一种有效的方法,将适配器插入文本到图像基础模型中,以实现复杂下游任务的同时保持基础模型的泛化能力。该方法的核心思想是优化与 2D 特征图相关的注意力机制,从而增强适配器的性能。该方法在 meme 视频生成任务上取得了显著成果。我们希望本工作为大型文本到图像模型的后训练任务提供思路。此外,由于该方法与 SD1.5 派生模型兼容性良好,对开源社区具有一定价值。因此,我们将发布相关代码(\url{https://songkey.github.io/hellomeme})。
引用
@article{arxiv.2410.22901,
title = {HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models},
author = {Shengkai Zhang and Nianhong Jiao and Tian Li and Chaojie Yang and Chenhui Xue and Boya Niu and Jun Gao},
journal= {arXiv preprint arXiv:2410.22901},
year = {2024}
}
备注
11 pages, 7 figures, 2 tables