English

Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners

Computer Vision and Pattern Recognition 2024-04-26 v3

Abstract

Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified in text prompts, can we leverage the powerful representations learned by pre-trained diffusion models for discriminative tasks such as image-text matching? To answer this question, we propose a novel approach, Discriminative Stable Diffusion (DSD), which turns pre-trained text-to-image diffusion models into few-shot discriminative learners. Our approach mainly uses the cross-attention score of a Stable Diffusion model to capture the mutual influence between visual and textual information and fine-tune the model via efficient attention-based prompt learning to perform image-text matching. By comparing DSD with state-of-the-art methods on several benchmark datasets, we demonstrate the potential of using pre-trained diffusion models for discriminative tasks with superior results on few-shot image-text matching.

Keywords

Cite

@article{arxiv.2305.10722,
  title  = {Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners},
  author = {Xuehai He and Weixi Feng and Tsu-Jui Fu and Varun Jampani and Arjun Akula and Pradyumna Narayana and Sugato Basu and William Yang Wang and Xin Eric Wang},
  journal= {arXiv preprint arXiv:2305.10722},
  year   = {2024}
}
R2 v1 2026-06-28T10:37:51.729Z