Contextual Procurement Auctions with Bandit Learning
Computer Science and Game Theory
2026-07-07 v1 Machine Learning
Abstract
We study repeated contextual procurement auctions in which the platform must learn context-dependent product values from bandit feedback. We give an exactly truthful explore-then-commit mechanism with regret. We also give a frozen-payment UCB mechanism with a regret-incentive tradeoff: the near-UCB tuning attains welfare regret, while for fixed its total incentive error is ; the balanced tuning gives on both scales. Regret is measured as welfare loss relative to the full-information efficient allocation. We prove a matching lower bound for the frozen-payment regret-incentive tradeoff.
Cite
@article{arxiv.2607.05813,
title = {Contextual Procurement Auctions with Bandit Learning},
author = {Yiling Chen and Shi Feng and Sadie Zhao},
journal= {arXiv preprint arXiv:2607.05813},
year = {2026}
}