English

Localized exploration in contextual dynamic pricing achieves dimension-free regret

Machine Learning 2024-12-30 v1 Machine Learning Optimization and Control

Abstract

We study the problem of contextual dynamic pricing with a linear demand model. We propose a novel localized exploration-then-commit (LetC) algorithm which starts with a pure exploration stage, followed by a refinement stage that explores near the learned optimal pricing policy, and finally enters a pure exploitation stage. The algorithm is shown to achieve a minimax optimal, dimension-free regret bound when the time horizon exceeds a polynomial of the covariate dimension. Furthermore, we provide a general theoretical framework that encompasses the entire time spectrum, demonstrating how to balance exploration and exploitation when the horizon is limited. The analysis is powered by a novel critical inequality that depicts the exploration-exploitation trade-off in dynamic pricing, mirroring its existing counterpart for the bias-variance trade-off in regularized regression. Our theoretical results are validated by extensive experiments on synthetic and real-world data.

Keywords

Cite

@article{arxiv.2412.19252,
  title  = {Localized exploration in contextual dynamic pricing achieves dimension-free regret},
  author = {Jinhang Chai and Yaqi Duan and Jianqing Fan and Kaizheng Wang},
  journal= {arXiv preprint arXiv:2412.19252},
  year   = {2024}
}

Comments

60 pages, 9 figures

R2 v1 2026-06-28T20:49:16.625Z