Super-Exponential Regret for UCT, AlphaGo and Variants
Machine Learning
2024-05-20 v2 Artificial Intelligence
Abstract
We improve the proofs of the lower bounds of Coquelin and Munos (2007) that demonstrate that UCT can have regret (with exp terms) on the -chain environment, and that a `polynomial' UCT variant has regret on the same environment -- the original proofs contain an oversight for rewards bounded in , which we fix in the present draft. We also adapt the proofs to AlphaGo's MCTS and its descendants (e.g., AlphaZero, Leela Zero) to also show regret.
Cite
@article{arxiv.2405.04407,
title = {Super-Exponential Regret for UCT, AlphaGo and Variants},
author = {Laurent Orseau and Remi Munos},
journal= {arXiv preprint arXiv:2405.04407},
year = {2024}
}