English

Vision-Language Models as a Source of Rewards

Machine Learning 2024-07-16 v3

Abstract

Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents. We show how rewards for visual achievement of a variety of language goals can be derived from the CLIP family of models, and used to train RL agents that can achieve a variety of language goals. We showcase this approach in two distinct visual domains and present a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.

Keywords

Cite

@article{arxiv.2312.09187,
  title  = {Vision-Language Models as a Source of Rewards},
  author = {Kate Baumli and Satinder Baveja and Feryal Behbahani and Harris Chan and Gheorghe Comanici and Sebastian Flennerhag and Maxime Gazeau and Kristian Holsheimer and Dan Horgan and Michael Laskin and Clare Lyle and Hussain Masoom and Kay McKinney and Volodymyr Mnih and Alexander Neitz and Dmitry Nikulin and Fabio Pardo and Jack Parker-Holder and John Quan and Tim Rocktäschel and Himanshu Sahni and Tom Schaul and Yannick Schroecker and Stephen Spencer and Richie Steigerwald and Luyu Wang and Lei Zhang},
  journal= {arXiv preprint arXiv:2312.09187},
  year   = {2024}
}

Comments

10 pages, 5 figures

R2 v1 2026-06-28T13:51:23.365Z