Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics
Abstract
The impact of a given training point on a statistical model is classically measured through its leave-one-out influence, which quantifies the effect of its removal from the training set on the model accuracy. While the statistics of leave-one-out influences are well understood in the low-dimensional, large sample limit , they become more intricate in high dimensions, as the influence of a given sample develops non-trivial dependencies on all other training samples. For convex M-estimation under Gaussian design, in the high-dimensional limit , we show that the distribution of the influences across the training set converges to a limiting measure which we sharply characterize. Building on these results, we provide evidence that influential samples tend to lie close to the decision boundary, thereby making contact with a standard data selection heuristic in active learning.
Cite
@article{arxiv.2607.09250,
title = {Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics},
author = {Hugo Cui},
journal= {arXiv preprint arXiv:2607.09250},
year = {2026}
}