English

Mixtures of Experts Unlock Parameter Scaling for Deep RL

Machine Learning 2024-06-27 v3 Artificial Intelligence

Abstract

The recent rapid progress in (self) supervised learning models is in large part predicted by empirical scaling laws: a model's performance scales proportionally to its size. Analogous scaling laws remain elusive for reinforcement learning domains, however, where increasing the parameter count of a model often hurts its final performance. In this paper, we demonstrate that incorporating Mixture-of-Expert (MoE) modules, and in particular Soft MoEs (Puigcerver et al., 2023), into value-based networks results in more parameter-scalable models, evidenced by substantial performance increases across a variety of training regimes and model sizes. This work thus provides strong empirical evidence towards developing scaling laws for reinforcement learning.

Keywords

Cite

@article{arxiv.2402.08609,
  title  = {Mixtures of Experts Unlock Parameter Scaling for Deep RL},
  author = {Johan Obando-Ceron and Ghada Sokar and Timon Willi and Clare Lyle and Jesse Farebrother and Jakob Foerster and Gintare Karolina Dziugaite and Doina Precup and Pablo Samuel Castro},
  journal= {arXiv preprint arXiv:2402.08609},
  year   = {2024}
}
R2 v1 2026-06-28T14:47:33.627Z