English

Parameter-free Optimal Rates for Nonlinear Semi-Norm Contractions with Applications to $Q$-Learning

Machine Learning 2026-03-24 v2

Abstract

Algorithms for solving \textit{nonlinear} fixed-point equations -- such as average-reward \textit{QQ-learning} and \textit{TD-learning} -- often involve semi-norm contractions. Achieving parameter-free optimal convergence rates for these methods via Polyak--Ruppert averaging has remained elusive, largely due to the non-monotonicity of such semi-norms. We close this gap by (i.) recasting the averaged error as a linear recursion involving a nonlinear perturbation, and (ii.) taming the nonlinearity by coupling the semi-norm's contraction with the monotonicity of a suitably induced norm. Our main result yields the first parameter-free O~(1/t)\tilde{O}(1/\sqrt{t}) optimal rates for QQ-learning in both average-reward and exponentially discounted settings, where tt denotes the iteration index. The result applies within a broad framework that accommodates synchronous and asynchronous updates, single-agent and distributed deployments, and data streams obtained either from simulators or along Markovian trajectories.

Keywords

Cite

@article{arxiv.2508.05984,
  title  = {Parameter-free Optimal Rates for Nonlinear Semi-Norm Contractions with Applications to $Q$-Learning},
  author = {Ankur Naskar and Gugan Thoppe and Vijay Gupta},
  journal= {arXiv preprint arXiv:2508.05984},
  year   = {2026}
}