English

Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs

Artificial Intelligence 2026-05-20 v1 Machine Learning Statistics Theory Machine Learning Statistics Theory

Abstract

We study reinforcement learning for episodic Markov Decision Processes (MDPs) whose transitions are modelled by a multinomial logistic (MNL) model. Existing algorithms for MNL mixture MDPs yield a regret of O~(dH2T)\smash{\tilde{O}(dH^2\sqrt{T})} (Li et al., 2024), where dd is the feature dimension, HH the episode length, and TT the number of episodes. Inspired by the logistic bandit literature (Abeille et al., 2021; Faury et al., 2022; Boudart et al., 2026), we introduce a problem-dependent constant σˉ_T1/2\bar\sigma\_T \leq 1/2, measuring the normalised average variance of the optimal downstream value function along the learner's trajectory. We propose an algorithm achieving a regret of O~(dH2σˉ_TT)\smash{\tilde{O}(dH^2\bar\sigma\_T\sqrt{T})}, which recovers the existing bound in the worst case and improves upon it for structured MDPs. For instance, for KL-constrained robust MDPs, σˉ_T=O(H1)\bar\sigma\_T = O(H^{-1}), reducing the horizon dependence by a factor HH. We further establish a matching Ω(dH2σˉ_TT)\smash{\Omega(dH^2\bar\sigma\_T\sqrt{T})} lower bound, proving minimax optimality (up to logarithmic factors) and fully characterising the regret complexity of MNL mixture MDPs for the first time.

Keywords

Cite

@article{arxiv.2605.19768,
  title  = {Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs},
  author = {Pierre Boudart and Pierre Gaillard and Alessandro Rudi},
  journal= {arXiv preprint arXiv:2605.19768},
  year   = {2026}
}
R2 v1 2026-07-22T07:21:38.901Z