English

Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning

Machine Learning 2026-05-18 v1 Machine Learning

Abstract

We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episode. Performance is measured by cumulative regret against the episode-wise optimal value, k=1K[V,MkVπk,Mk]\sum_{k=1}^K [V^{*,M^k} - V^{\pi^k,M^k}], where MkM^k represents the action context in the kk-th episode. We show that the MVP algorithm naturally extends to this framework and enjoys strong theoretical guarantees. In particular, we establish a minimax regret bound of O~(SAH3KlogL)\widetilde{O}(\sqrt{SAH^3K\log L}) for adversarial contexts, where LL denotes the number of possible contexts. This result implies a regret bound of O~(SAH3K)\widetilde{O}(\sqrt{SAH^3K}) for stochastic contexts. We further translate the stochastic regret guarantee into a sample complexity bound of O~(SAH3/ϵ2)\widetilde{O}(SAH^3/\epsilon^2) for a fixed context distribution. In addition, we derive a gap-dependent regret bound of O~(infp[0,1)(1Δminp+pKΔminp)logKpoly(S,A,H)), \widetilde O\left( \inf_{p\in [0,1)} \left( \frac{1}{\Delta_{\min}^{p}} + pK\Delta_{\min}^{p} \right)\log K \cdot \mathrm{poly}(S,A,H) \right), where Δminp\Delta_{\min}^{p} is the global pp-trimmed positive-gap floor over suboptimal (h,s,a)(h,s,a) triples. This bound can substantially improve upon the minimax rate when the relevant suboptimality gaps are large.

Keywords

Cite

@article{arxiv.2605.15692,
  title  = {Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning},
  author = {Zijun Chen and Zihan Zhang},
  journal= {arXiv preprint arXiv:2605.15692},
  year   = {2026}
}
R2 v1 2026-07-22T07:13:52.273Z