Analysis of Search Heuristics in the Multi-Armed Bandit Setting
Abstract
We consider the classic Multi-Armed Bandit setting to understand the exploration/exploitation tradeoffs made by different search heuristics. Since many search heuristics work by comparing different options (in evolutionary algorithms called "individuals"; in the Bandit literature called "arms"), we work with the "Dueling Bandits" setting. In each iteration, a comparison between different arms can be made; in the binary stochastic setting, each arm has a fixed winning probability against any other arm. A Condorcet winner is any arm that beats every other arm with a probability strictly higher than . We show that evolutionary algorithms are rather bad at identifying the Condorcet winner: Even if the Condorcet winner beats every other arm with a probability , the (1+1) EA, in its stationary distribution, chooses the Condorcet winner only with constant probability if . By contrast, we show that a simple EDA (based on the Max-Min Ant System with iteration-best update) will choose the Condorcet winner in its maintained distribution with probability . As a remedy for the (1+1) EA, we show how repeated duels can significantly boost the probability of the Condorcet winner in the stationary distribution.
Keywords
Cite
@article{arxiv.2604.08109,
title = {Analysis of Search Heuristics in the Multi-Armed Bandit Setting},
author = {Jasmin Brandt and Barbara Hammer and Timo Kötzing and Jurek Sander},
journal= {arXiv preprint arXiv:2604.08109},
year = {2026}
}
Comments
16 pages, 5 figures, GECCO 2026