English

SGDiff: A Style Guided Diffusion Model for Fashion Synthesis

Computer Vision and Pattern Recognition 2023-08-16 v1 Artificial Intelligence Multimedia

Abstract

This paper reports on the development of \textbf{a novel style guided diffusion model (SGDiff)} which overcomes certain weaknesses inherent in existing models for image synthesis. The proposed SGDiff combines image modality with a pretrained text-to-image diffusion model to facilitate creative fashion image synthesis. It addresses the limitations of text-to-image diffusion models by incorporating supplementary style guidance, substantially reducing training costs, and overcoming the difficulties of controlling synthesized styles with text-only inputs. This paper also introduces a new dataset -- SG-Fashion, specifically designed for fashion image synthesis applications, offering high-resolution images and an extensive range of garment categories. By means of comprehensive ablation study, we examine the application of classifier-free guidance to a variety of conditions and validate the effectiveness of the proposed model for generating fashion images of the desired categories, product attributes, and styles. The contributions of this paper include a novel classifier-free guidance method for multi-modal feature fusion, a comprehensive dataset for fashion image synthesis application, a thorough investigation on conditioned text-to-image synthesis, and valuable insights for future research in the text-to-image synthesis domain. The code and dataset are available at: \url{https://github.com/taited/SGDiff}.

Keywords

Cite

@article{arxiv.2308.07605,
  title  = {SGDiff: A Style Guided Diffusion Model for Fashion Synthesis},
  author = {Zhengwentai Sun and Yanghong Zhou and Honghong He and P. Y. Mok},
  journal= {arXiv preprint arXiv:2308.07605},
  year   = {2023}
}

Comments

Accepted by ACM MM'23

R2 v1 2026-06-28T11:55:49.368Z