Learning in Prophet Inequalities with Noisy Observations
Abstract
We study the prophet inequality, a fundamental problem in online decision-making and optimal stopping, in a practical setting where rewards are observed only through noisy realizations and reward distributions are unknown. At each stage, the decision-maker receives a noisy reward whose true value follows a linear model with an unknown latent parameter, and observes a feature vector drawn from a distribution. To address this challenge, we propose algorithms that integrate learning and decision-making via lower-confidence-bound (LCB) thresholding. In the i.i.d.\ setting, we establish that both an Explore-then-Decide strategy and an -Greedy variant achieve the sharp competitive ratio of , under a mild condition on the optimal value. For non-identical distributions, we show that a competitive ratio of can be guaranteed against a relaxed benchmark. Moreover, with limited window access to past rewards, the tight ratio of against the optimal benchmark is achieved.
Cite
@article{arxiv.2604.01789,
title = {Learning in Prophet Inequalities with Noisy Observations},
author = {Jung-hun Kim and Vianney Perchet},
journal= {arXiv preprint arXiv:2604.01789},
year = {2026}
}
Comments
ICLR 2026