English

Online Action Learning in High Dimensions: A Conservative Perspective

Machine Learning 2024-03-26 v4 Machine Learning Econometrics

Abstract

Sequential learning problems are common in several fields of research and practical applications. Examples include dynamic pricing and assortment, design of auctions and incentives and permeate a large number of sequential treatment experiments. In this paper, we extend one of the most popular learning solutions, the ϵt\epsilon_t-greedy heuristics, to high-dimensional contexts considering a conservative directive. We do this by allocating part of the time the original rule uses to adopt completely new actions to a more focused search in a restrictive set of promising actions. The resulting rule might be useful for practical applications that still values surprises, although at a decreasing rate, while also has restrictions on the adoption of unusual actions. With high probability, we find reasonable bounds for the cumulative regret of a conservative high-dimensional decaying ϵt\epsilon_t-greedy rule. Also, we provide a lower bound for the cardinality of the set of viable actions that implies in an improved regret bound for the conservative version when compared to its non-conservative counterpart. Additionally, we show that end-users have sufficient flexibility when establishing how much safety they want, since it can be tuned without impacting theoretical properties. We illustrate our proposal both in a simulation exercise and using a real dataset.

Keywords

Cite

@article{arxiv.2009.13961,
  title  = {Online Action Learning in High Dimensions: A Conservative Perspective},
  author = {Claudio Cardoso Flores and Marcelo Cunha Medeiros},
  journal= {arXiv preprint arXiv:2009.13961},
  year   = {2024}
}

Comments

We found an error in the proof of the main theorem which cannot be fixed without completely changing the results in the paper

R2 v1 2026-06-23T18:52:37.144Z