English

A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning

Machine Learning 2026-05-19 v1 Artificial Intelligence

Abstract

We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. Across poker variants, matrix games, a dice game, and multiple learning algorithms, eliminating all positive-reach contingent decisions causes rapid convergence to a deterministic exploitation attractor, a fixed point at near-maximal loss. Preserving even a single positive-reach contingent decision point prevents this collapse. A frozen baseline and fixed-opponent control confirm that the mechanism is co-adaptation under constraint, not the perturbation itself. The phenomenon is timing-invariant, fully reversible upon action restoration, and intensifies under function approximation. These results establish a sharp threshold at zero reach-weighted contingent action capacity, with severity scaling continuously via reach-weighted capacity in the tested domains.

Keywords

Cite

@article{arxiv.2605.16315,
  title  = {A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning},
  author = {Arahan Kujur},
  journal= {arXiv preprint arXiv:2605.16315},
  year   = {2026}
}

Comments

18 pages, 7 figures

R2 v1 2026-07-22T07:15:14.202Z