English

Analysis of Search Heuristics in the Multi-Armed Bandit Setting

Neural and Evolutionary Computing 2026-04-10 v1

Abstract

We consider the classic Multi-Armed Bandit setting to understand the exploration/exploitation tradeoffs made by different search heuristics. Since many search heuristics work by comparing different options (in evolutionary algorithms called "individuals"; in the Bandit literature called "arms"), we work with the "Dueling Bandits" setting. In each iteration, a comparison between different arms can be made; in the binary stochastic setting, each arm has a fixed winning probability against any other arm. A Condorcet winner is any arm that beats every other arm with a probability strictly higher than 1/21/2. We show that evolutionary algorithms are rather bad at identifying the Condorcet winner: Even if the Condorcet winner beats every other arm with a probability 1p1-p, the (1+1) EA, in its stationary distribution, chooses the Condorcet winner only with constant probability if p=Ω(1/n)p=\Omega(1/n). By contrast, we show that a simple EDA (based on the Max-Min Ant System with iteration-best update) will choose the Condorcet winner in its maintained distribution with probability 1Θ(p)1-\Theta(p). As a remedy for the (1+1) EA, we show how repeated duels can significantly boost the probability of the Condorcet winner in the stationary distribution.

Keywords

Cite

@article{arxiv.2604.08109,
  title  = {Analysis of Search Heuristics in the Multi-Armed Bandit Setting},
  author = {Jasmin Brandt and Barbara Hammer and Timo Kötzing and Jurek Sander},
  journal= {arXiv preprint arXiv:2604.08109},
  year   = {2026}
}

Comments

16 pages, 5 figures, GECCO 2026