English

VLRM: Vision-Language Models act as Reward Models for Image Captioning

Computer Vision and Pattern Recognition 2024-04-03 v1

Abstract

In this work, we present an unsupervised method for enhancing an image captioning model (in our case, BLIP2) using reinforcement learning and vision-language models like CLIP and BLIP2-ITM as reward models. The RL-tuned model is able to generate longer and more comprehensive descriptions. Our model reaches impressive 0.90 R@1 CLIP Recall score on MS-COCO Carpathy Test Split. Weights are available at https://huggingface.co/sashakunitsyn/vlrm-blip2-opt-2.7b.

Keywords

Cite

@article{arxiv.2404.01911,
  title  = {VLRM: Vision-Language Models act as Reward Models for Image Captioning},
  author = {Maksim Dzabraev and Alexander Kunitsyn and Andrei Ivaniuta},
  journal= {arXiv preprint arXiv:2404.01911},
  year   = {2024}
}
R2 v1 2026-06-28T15:41:38.103Z