English

Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

Machine Learning 2026-07-09 v1 Machine Learning

Abstract

In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point ηm\eta_m induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator ηm(n)\eta_m^{(n)} based on an empirical Markov decision process. For a fixed number of quantiles mm, we establish a non-asymptotic error bound for ηm(n)\eta_m^{(n)} and ηm\eta_m under the supremum WW_\infty metric, showing that the estimation error scales as O~(m/n)\widetilde{O}(\sqrt{m/n}) with respect to mm and nn. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric n\sqrt{n} convergence rate. We derive the asymptotic distribution of the quantile parameters n(θm(n)θm)\sqrt{n}(\theta_m^{(n)}-\theta_m) and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals n(ηm(n)(s)ηm(s))f\sqrt{n}(\eta_m^{(n)}(s)-\eta_m(s))f, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.

Cite

@article{arxiv.2607.08444,
  title  = {Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning},
  author = {Zijie Cheng and Yang Peng and Zhihua Zhang},
  journal= {arXiv preprint arXiv:2607.08444},
  year   = {2026}
}