English

Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Machine Learning 2019-04-26 v1 Artificial Intelligence Machine Learning

Abstract

Rather than proposing a new method, this paper investigates an issue present in existing learning algorithms. We study the learning dynamics of reinforcement learning (RL), specifically a characteristic coupling between learning and data generation that arises because RL agents control their future data distribution. In the presence of function approximation, this coupling can lead to a problematic type of 'ray interference', characterized by learning dynamics that sequentially traverse a number of performance plateaus, effectively constraining the agent to learn one thing at a time even when learning in parallel is better. We establish the conditions under which ray interference occurs, show its relation to saddle points and obtain the exact learning dynamics in a restricted setting. We characterize a number of its properties and discuss possible remedies.

Keywords

Cite

@article{arxiv.1904.11455,
  title  = {Ray Interference: a Source of Plateaus in Deep Reinforcement Learning},
  author = {Tom Schaul and Diana Borsa and Joseph Modayil and Razvan Pascanu},
  journal= {arXiv preprint arXiv:1904.11455},
  year   = {2019}
}

Comments

Full version of RLDM abstract

R2 v1 2026-06-23T08:49:37.416Z