English

More Control for Free! Image Synthesis with Semantic Diffusion Guidance

Computer Vision and Pattern Recognition 2022-12-06 v4 Graphics

Abstract

Controllable image synthesis models allow creation of diverse images based on text instructions or guidance from a reference image. Recently, denoising diffusion probabilistic models have been shown to generate more realistic imagery than prior methods, and have been successfully demonstrated in unconditional and class-conditional settings. We investigate fine-grained, continuous control of this model class, and introduce a novel unified framework for semantic diffusion guidance, which allows either language or image guidance, or both. Guidance is injected into a pretrained unconditional diffusion model using the gradient of image-text or image matching scores, without re-training the diffusion model. We explore CLIP-based language guidance as well as both content and style-based image guidance in a unified framework. Our text-guided synthesis approach can be applied to datasets without associated text annotations. We conduct experiments on FFHQ and LSUN datasets, and show results on fine-grained text-guided image synthesis, synthesis of images related to a style or content reference image, and examples with both textual and image guidance.

Keywords

Cite

@article{arxiv.2112.05744,
  title  = {More Control for Free! Image Synthesis with Semantic Diffusion Guidance},
  author = {Xihui Liu and Dong Huk Park and Samaneh Azadi and Gong Zhang and Arman Chopikyan and Yuxiao Hu and Humphrey Shi and Anna Rohrbach and Trevor Darrell},
  journal= {arXiv preprint arXiv:2112.05744},
  year   = {2022}
}

Comments

WACV 2023. Project page https://xh-liu.github.io/sdg/

R2 v1 2026-06-24T08:12:46.727Z