English

Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning

Machine Learning 2024-05-24 v2 Information Theory Systems and Control Systems and Control math.IT Statistics Theory Machine Learning Statistics Theory

Abstract

This paper studies reward-agnostic exploration in reinforcement learning (RL) -- a scenario where the learner is unware of the reward functions during the exploration stage -- and designs an algorithm that improves over the state of the art. More precisely, consider a finite-horizon inhomogeneous Markov decision process with SS states, AA actions, and horizon length HH, and suppose that there are no more than a polynomial number of given reward functions of interest. By collecting an order of \begin{align*} \frac{SAH^3}{\varepsilon^2} \text{ sample episodes (up to log factor)} \end{align*} without guidance of the reward information, our algorithm is able to find ε\varepsilon-optimal policies for all these reward functions, provided that ε\varepsilon is sufficiently small. This forms the first reward-agnostic exploration scheme in this context that achieves provable minimax optimality. Furthermore, once the sample size exceeds S2AH3ε2\frac{S^2AH^3}{\varepsilon^2} episodes (up to log factor), our algorithm is able to yield ε\varepsilon accuracy for arbitrarily many reward functions (even when they are adversarially designed), a task commonly dubbed as ``reward-free exploration.'' The novelty of our algorithm design draws on insights from offline RL: the exploration scheme attempts to maximize a critical reward-agnostic quantity that dictates the performance of offline RL, while the policy learning paradigm leverages ideas from sample-optimal offline RL paradigms.

Keywords

Cite

@article{arxiv.2304.07278,
  title  = {Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning},
  author = {Gen Li and Yuling Yan and Yuxin Chen and Jianqing Fan},
  journal= {arXiv preprint arXiv:2304.07278},
  year   = {2024}
}

Comments

accepted for presentation in COLT 2024