Strategic Multi-Armed Bandit Problems Under Debt-Free Reporting
Abstract
We consider the classical multi-armed bandit problem, but with strategic arms. In this context, each arm is characterized by a bounded support reward distribution and strategically aims to maximize its own utility by potentially retaining a portion of its reward, and disclosing only a fraction of it to the learning agent. This scenario unfolds as a game over rounds, leading to a competition of objectives between the learning agent, aiming to minimize their regret, and the arms, motivated by the desire to maximize their individual utilities. To address these dynamics, we introduce a new mechanism that establishes an equilibrium wherein each arm behaves truthfully and discloses as much of its rewards as possible. With this mechanism, the agent can attain the second-highest average (true) reward among arms, with a cumulative regret bounded by (problem-dependent) or (worst-case).
Cite
@article{arxiv.2501.16018,
title = {Strategic Multi-Armed Bandit Problems Under Debt-Free Reporting},
author = {Ahmed Ben Yahmed and Clément Calauzènes and Vianney Perchet},
journal= {arXiv preprint arXiv:2501.16018},
year = {2025}
}