English

Foundation Models for Semantic Novelty in Reinforcement Learning

Machine Learning 2022-11-10 v1 Artificial Intelligence

Abstract

Effectively exploring the environment is a key challenge in reinforcement learning (RL). We address this challenge by defining a novel intrinsic reward based on a foundation model, such as contrastive language image pretraining (CLIP), which can encode a wealth of domain-independent semantic visual-language knowledge about the world. Specifically, our intrinsic reward is defined based on pre-trained CLIP embeddings without any fine-tuning or learning on the target RL task. We demonstrate that CLIP-based intrinsic rewards can drive exploration towards semantically meaningful states and outperform state-of-the-art methods in challenging sparse-reward procedurally-generated environments.

Keywords

Cite

@article{arxiv.2211.04878,
  title  = {Foundation Models for Semantic Novelty in Reinforcement Learning},
  author = {Tarun Gupta and Peter Karkus and Tong Che and Danfei Xu and Marco Pavone},
  journal= {arXiv preprint arXiv:2211.04878},
  year   = {2022}
}

Comments

Foundation Models for Decision Making Workshop at Neural Information Processing Systems, 2022

R2 v1 2026-06-28T05:30:51.922Z