English

Conformal-Style Quantile Analyses for Stochastic Bandits

Machine Learning 2026-05-11 v1 Machine Learning

Abstract

Stochastic bandit algorithms are usually analyzed under a mean-reward criterion, yet many problems favor arms with strong upper-tail performance, which we study herein. For a fixed miscoverage level α\alpha, the natural upper-tail target of arm jj is the upper endpoint Fj1(1α/2)F_j^{-1}(1-\alpha/2) of a central prediction interval. This target can rank arms differently from their means, creating a central mismatch with the classical bandit objective. To this end, we propose ACP-UCB1, a conformal-style policy that combines an adaptive conformal estimate of the upper endpoint with a UCB-type optimism bonus. The technical challenge is that the conformity scores used by ACP-UCB1 are recomputed from evolving empirical quantile estimates and evaluated at an adaptive level. We control this endpoint through reward-quantile concentration, a perturbation argument for recomputed score quantiles, and deterministic localization of the adaptive level. ACP-UCB1 achieves logarithmic upper-quantile regret with per-arm contribution O(\nicefraclognΔjACP)O(\nicefrac{\log n}{\Delta_j^{\mathrm{ACP}}}). We also provide metric-specific regret decompositions comparing ACP-UCB1 with UCB1 and use numerical experiments to validate performance and improvement.

Keywords

Cite

@article{arxiv.2605.07115,
  title  = {Conformal-Style Quantile Analyses for Stochastic Bandits},
  author = {Chengyu Du and Mengfan Xu},
  journal= {arXiv preprint arXiv:2605.07115},
  year   = {2026}
}
R2 v1 2026-07-01T12:56:40.989Z