The Structure of Cross-Validation Error: Stability, Covariance, and Minimax Limits
Abstract
Despite ongoing theoretical research on cross-validation (CV), many theoretical questions remain widely open. This motivates our investigation into how properties of algorithm-distribution pairs can affect the choice for the number of folds in -fold CV. Our results consist of a novel decomposition of the mean-squared error of cross-validation for risk estimation, which explicitly captures the correlations of error estimates across overlapping folds and includes a novel algorithmic stability notion, squared loss stability, that is considerably weaker than the typically required hypothesis stability in other comparable works. Furthermore, we prove: 1. For any learning algorithm that minimizes empirical risk, the mean-squared error of the -fold cross-validation estimator of the population risk satisfies the following minimax lower bound: where is the sample size, the number of folds, and denotes the number of folds attaining the minimax optimum. This shows that even under idealized conditions, for large values of , CV cannot attain the optimum of order achievable by a validation set of size , reflecting an inherent penalty caused by dependence between folds. 2. Complementing this, we exhibit learning rules for which matching (up to constants) the accuracy of a hold-out estimator of a single fold of size . Together these results delineate the fundamental trade-off in resampling-based risk estimation: CV cannot fully exploit all samples for unbiased risk evaluation, and its minimax performance is pinned between the and regimes.
Keywords
Cite
@article{arxiv.2511.03554,
title = {The Structure of Cross-Validation Error: Stability, Covariance, and Minimax Limits},
author = {Ido Nachum and Rüdiger Urbanke and Thomas Weinberger},
journal= {arXiv preprint arXiv:2511.03554},
year = {2026}
}
Comments
60 pages