English

A Selection Premium Decomposition for the Expected Maximum of Random Walks

Statistics Theory 2026-02-24 v1 Statistics Theory

Abstract

When KK models are evaluated on the same validation set of size nn, the selected winner's apparent performance is biased upward. Suppose KK models are evaluated on a shared sequence of i.i.d. observations X1,,XnX_1,\dots, X_n, where model kk achieves response fk(Xi)f_k(X_i) with mean μk=E[fk(X)]\mu_k = \mathbb E[f_k(X)]. Writing Yi,k=fk(Xi)μkY_{i,k} = f_k(X_i)-\mu_k for the centered increment and Sn,k=i=1nYi,kS_{n,k} = \sum_{i=1}^n Y_{i,k} for the centered cumulative score, the expected maximum satisfies 0E[maxkSn,k]=i=1nE[φK(Si1)]0\le\mathbb E\bigl[\max_k S_{n,k}\bigr] = \sum_{i=1}^n \mathbb E\bigl[\varphi_K(S_{i-1})\bigr] where φK(u)=E[maxk(uk+Yk)]maxkuk\varphi_K(u) = \mathbb{E}\bigl[\max_k(u_k + Y_k)\bigr] - \max_k u_k, uRKu\in \mathbb R^K, is the selection premium function. This formula corresponds to the null hypothesis case (all models are equal in the sense that they have the same mean), which clarifies that the bias arises from selection. While this decomposition follows from elementary conditioning and telescoping, we develop the analytical consequences in five directions. (i) structural properties of φK\varphi_K; (ii) extension to stopping times, recovering Wald's equation at K=1K=1; (iii) a winner's curse decomposition for heterogeneous means; (iv) a universal bias concentration law showing that the first α\alpha-fraction of observations generates a α\sqrt\alpha-fraction of total bias.

Keywords

Cite

@article{arxiv.2602.19481,
  title  = {A Selection Premium Decomposition for the Expected Maximum of Random Walks},
  author = {Victor H. de la Pena and Fangyuan Lin and Victor K. de la Pena},
  journal= {arXiv preprint arXiv:2602.19481},
  year   = {2026}
}
R2 v1 2026-07-01T10:46:50.047Z