English

VO$Q$L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

Machine Learning 2022-12-13 v1 Machine Learning

Abstract

We study time-inhomogeneous episodic reinforcement learning (RL) under general function approximation and sparse rewards. We design a new algorithm, Variance-weighted Optimistic QQ-Learning (VOQQL), based on QQ-learning and bound its regret assuming completeness and bounded Eluder dimension for the regression function class. As a special case, VOQQL achieves O~(dHT+d6H5)\tilde{O}(d\sqrt{HT}+d^6H^{5}) regret over TT episodes for a horizon HH MDP under (dd-dimensional) linear function approximation, which is asymptotically optimal. Our algorithm incorporates weighted regression-based upper and lower bounds on the optimal value function to obtain this improved regret. The algorithm is computationally efficient given a regression oracle over the function class, making this the first computationally tractable and statistically optimal approach for linear MDPs.

Keywords

Cite

@article{arxiv.2212.06069,
  title  = {VO$Q$L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation},
  author = {Alekh Agarwal and Yujia Jin and Tong Zhang},
  journal= {arXiv preprint arXiv:2212.06069},
  year   = {2022}
}
R2 v1 2026-06-28T07:31:29.558Z