English

HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior

Computer Vision and Pattern Recognition 2024-12-02 v1

Abstract

Text-to-image diffusion models have emerged as powerful priors for real-world image super-resolution (Real-ISR). However, existing methods may produce unintended results due to noisy text prompts and their lack of spatial information. In this paper, we present HoliSDiP, a framework that leverages semantic segmentation to provide both precise textual and spatial guidance for diffusion-based Real-ISR. Our method employs semantic labels as concise text prompts while introducing dense semantic guidance through segmentation masks and our proposed Segmentation-CLIP Map. Extensive experiments demonstrate that HoliSDiP achieves significant improvement in image quality across various Real-ISR scenarios through reduced prompt noise and enhanced spatial control.

Keywords

Cite

@article{arxiv.2411.18662,
  title  = {HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior},
  author = {Li-Yuan Tsao and Hao-Wei Chen and Hao-Wei Chung and Deqing Sun and Chun-Yi Lee and Kelvin C. K. Chan and Ming-Hsuan Yang},
  journal= {arXiv preprint arXiv:2411.18662},
  year   = {2024}
}

Comments

Project page: https://liyuantsao.github.io/HoliSDiP/

R2 v1 2026-06-28T20:15:06.192Z