Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
Abstract
We study reinforcement learning for episodic Markov Decision Processes (MDPs) whose transitions are modelled by a multinomial logistic (MNL) model. Existing algorithms for MNL mixture MDPs yield a regret of (Li et al., 2024), where is the feature dimension, the episode length, and the number of episodes. Inspired by the logistic bandit literature (Abeille et al., 2021; Faury et al., 2022; Boudart et al., 2026), we introduce a problem-dependent constant , measuring the normalised average variance of the optimal downstream value function along the learner's trajectory. We propose an algorithm achieving a regret of , which recovers the existing bound in the worst case and improves upon it for structured MDPs. For instance, for KL-constrained robust MDPs, , reducing the horizon dependence by a factor . We further establish a matching lower bound, proving minimax optimality (up to logarithmic factors) and fully characterising the regret complexity of MNL mixture MDPs for the first time.
Cite
@article{arxiv.2605.19768,
title = {Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs},
author = {Pierre Boudart and Pierre Gaillard and Alessandro Rudi},
journal= {arXiv preprint arXiv:2605.19768},
year = {2026}
}