English

QGFN: Controllable Greediness with Action Values

Machine Learning 2024-11-04 v3

Abstract

Generative Flow Networks (GFlowNets; GFNs) are a family of energy-based generative methods for combinatorial objects, capable of generating diverse and high-utility samples. However, consistently biasing GFNs towards producing high-utility samples is non-trivial. In this work, we leverage connections between GFNs and reinforcement learning (RL) and propose to combine the GFN policy with an action-value estimate, QQ, to create greedier sampling policies which can be controlled by a mixing parameter. We show that several variants of the proposed method, QGFN, are able to improve on the number of high-reward samples generated in a variety of tasks without sacrificing diversity.

Keywords

Cite

@article{arxiv.2402.05234,
  title  = {QGFN: Controllable Greediness with Action Values},
  author = {Elaine Lau and Stephen Zhewen Lu and Ling Pan and Doina Precup and Emmanuel Bengio},
  journal= {arXiv preprint arXiv:2402.05234},
  year   = {2024}
}

Comments

Accepted by 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

R2 v1 2026-06-28T14:42:13.462Z