English

Contextual Procurement Auctions with Bandit Learning

Computer Science and Game Theory 2026-07-07 v1 Machine Learning

Abstract

We study repeated contextual procurement auctions in which the platform must learn context-dependent product values from bandit feedback. We give an exactly truthful explore-then-commit mechanism with O~((ng)1/3T2/3)\widetilde O((ng)^{1/3}T^{2/3}) regret. We also give a frozen-payment UCB mechanism with a regret-incentive tradeoff: the near-UCB tuning attains O~(ngT)\widetilde O(\sqrt{ngT}) welfare regret, while for fixed n,gn,g its total incentive error is O~(T3/4)\widetilde O(T^{3/4}); the balanced tuning gives O~(T2/3)\widetilde O(T^{2/3}) on both scales. Regret is measured as welfare loss relative to the full-information efficient allocation. We prove a matching lower bound for the frozen-payment regret-incentive tradeoff.

Cite

@article{arxiv.2607.05813,
  title  = {Contextual Procurement Auctions with Bandit Learning},
  author = {Yiling Chen and Shi Feng and Sadie Zhao},
  journal= {arXiv preprint arXiv:2607.05813},
  year   = {2026}
}