中文

SEGA:利用语义引导指导文本到图像模型

计算机视觉与模式识别 2023-11-06 v2 人工智能 机器学习

摘要

文本到图像扩散模型近来因其仅从文本生成高保真图像的惊人能力而备受关注。然而,实现与用户意图一致的一次性生成几乎不可能,且对输入提示的微小改动常导致截然不同的图像。这使用户几乎无法进行控制语义。为了让用户掌控,我们展示了如何与扩散过程交互以沿语义方向灵活引导它。该语义引导(SEGA)可推广到任何使用无分类器引导的生成架构。更重要的是,它允许细微与大幅编辑、构图与风格变化,以及整体艺术构思的优化。我们在 Stable Diffusion、Paella 和 DeepFloyd-IF 等潜空间与基于像素的扩散模型上利用多种任务展示了 SEGA 的有效性,从而为其通用性、灵活性及对现有方法的改进提供了有力证据。

关键词

引用

@article{arxiv.2301.12247,
  title  = {SEGA: Instructing Text-to-Image Models using Semantic Guidance},
  author = {Manuel Brack and Felix Friedrich and Dominik Hintersdorf and Lukas Struppek and Patrick Schramowski and Kristian Kersting},
  journal= {arXiv preprint arXiv:2301.12247},
  year   = {2023}
}

备注

arXiv admin note: text overlap with arXiv:2212.06013 Proceedings of the Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS)