English

Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity

Computer Vision and Pattern Recognition 2024-05-24 v2 Machine Learning

Abstract

Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image compression algorithm that achieves both high perceptual and pixel-wise fidelity. In particular, we propose a compression framework that leverages text information mainly by text-adaptive encoding and training with joint image-text loss. By doing so, we avoid decoding based on text-guided generative models -- known for high generative diversity -- and effectively utilize the semantic information of text at a global level. Experimental results on various datasets show that our method can achieve high pixel-level and perceptual quality, with either human- or machine-generated captions. In particular, our method outperforms all baselines in terms of LPIPS, with some room for even more improvements when we use more carefully generated captions.

Keywords

Cite

@article{arxiv.2403.02944,
  title  = {Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity},
  author = {Hagyeong Lee and Minkyu Kim and Jun-Hyuk Kim and Seungeon Kim and Dokwan Oh and Jaeho Lee},
  journal= {arXiv preprint arXiv:2403.02944},
  year   = {2024}
}

Comments

The first two authors contributed equally

R2 v1 2026-06-28T15:09:45.595Z