English
Related papers

Related papers: Sharp concentration of uniform generalization erro…

200 papers

We consider diffraction at random point scatterers on general discrete point sets in $\R^\nu$, restricted to a finite volume. We allow for random amplitudes and random dislocations of the scatterers. We investigate the speed of convergence…

Mathematical Physics · Physics 2007-05-23 C. Kuelske

We prove an isoperimetric inequality for the uniform measure on a uniformly convex body and for a class of uniformly log-concave measures (that we introduce). These inequalities imply (up to universal constants) the log-Sobolev inequalities…

Probability · Mathematics 2008-02-01 Emanuel Milman , Sasha Sodin

We present a new PAC-Bayesian generalization bound. Standard bounds contain a $\sqrt{L_n \cdot \KL/n}$ complexity term which dominates unless $L_n$, the empirical error of the learning algorithm's randomized predictions, vanishes. We manage…

Machine Learning · Computer Science 2021-12-16 Zakaria Mhammedi , Peter D. Grunwald , Benjamin Guedj

Poincar{\'e} inequalities are ubiquitous in probability and analysis and have various applications in statistics (concentration of measure, rate of convergence of Markov chains). The Poincar{\'e} constant, for which the inequality is tight,…

Probability · Mathematics 2019-11-25 Loucas Pillaud-Vivien , Francis Bach , Tony Lelièvre , Alessandro Rudi , Gabriel Stoltz

We investigate the asymptotic behavior of Bayesian posterior distributions under independent and identically distributed ($i.i.d.$) misspecified models. More specifically, we study the concentration of the posterior distribution on…

Statistics Theory · Mathematics 2015-12-04 R. V. Ramamoorthi , Karthik Sriram , Ryan Martin

This article studies the achievable guarantees on the error rates of certain learning algorithms, with particular focus on refining logarithmic factors. Many of the results are based on a general technique for obtaining bounds on the error…

Machine Learning · Computer Science 2016-09-13 Steve Hanneke

In this paper we prove multilevel concentration inequalities for bounded functionals $f = f(X_1, \ldots, X_n)$ of random variables $X_1, \ldots, X_n$ that are either independent or satisfy certain logarithmic Sobolev inequalities. The…

Probability · Mathematics 2020-06-16 Friedrich Götze , Holger Sambale , Arthur Sinulis

In this paper we establish some explicit and sharp estimates of the spectral gap and the log-Sobolev constant for mean field particles system, uniform in the number of particles, when the confinement potential have many local minimums. Our…

Probability · Mathematics 2019-09-17 Arnaud Guillin , Wei Liu , Liming Wu , Chaoen Zhang

We study problem-dependent rates, i.e., generalization errors that scale near-optimally with the variance, the effective loss, or the gradient norms evaluated at the "best hypothesis." We introduce a principled framework dubbed "uniform…

Machine Learning · Statistics 2020-12-25 Yunbei Xu , Assaf Zeevi

When data contains measurement errors, it is necessary to make assumptions relating the observed, erroneous data to the unobserved true phenomena of interest. These assumptions should be justifiable on substantive grounds, but are often…

Machine Learning · Statistics 2020-12-24 Noam Finkelstein , Roy Adams , Suchi Saria , Ilya Shpitser

We investigate concentration properties of functions of random vectors with values in the discrete cube, satisfying the stochastic covering property (SCP) or the strong Rayleigh property (SRP). Our result for SCP measures include…

Probability · Mathematics 2021-08-31 Radosław Adamczak , Bartłomiej Polaczyk

In this article, we study rates of convergence of the generalization error of multi-class margin classifiers. In particular, we develop an upper bound theory quantifying the generalization error of various large margin classifiers. The…

Statistics Theory · Mathematics 2011-11-10 Xiaotong Shen , Lifeng Wang

In this paper, we study the convergence of the spectral embeddings obtained from the leading eigenvectors of certain similarity matrices to their population counterparts. We opt to study this convergence in a uniform (instead of average)…

Statistics Theory · Mathematics 2023-04-26 Ruofei Zhao , Songkai Xue , Yuekai Sun

While the expected calibration error (ECE), which employs binning, is widely adopted to evaluate the calibration performance of machine learning models, theoretical understanding of its estimation bias is limited. In this paper, we present…

Machine Learning · Computer Science 2025-05-27 Futoshi Futami , Masahiro Fujisawa

Deep neural networks generalize well despite being exceedingly overparameterized and being trained without explicit regularization. This curious phenomenon has inspired extensive research activity in establishing its statistical principles:…

Machine Learning · Statistics 2021-09-16 Ke Wang , Christos Thrampoulidis

We study the problem of generalized uniformity testing \cite{BC17} of a discrete probability distribution: Given samples from a probability distribution $p$ over an {\em unknown} discrete domain $\mathbf{\Omega}$, we want to distinguish,…

Data Structures and Algorithms · Computer Science 2017-09-08 Ilias Diakonikolas , Daniel M. Kane , Alistair Stewart

We extend recent higher order concentration results in the discrete setting to include functions of possibly dependent variables whose distribution (on the product space) satisfies a logarithmic Sobolev inequality with respect to a…

Probability · Mathematics 2020-05-15 Friedrich Götze , Holger Sambale , Arthur Sinulis

The vast majority of statistical theory on binary classification characterizes performance in terms of accuracy. However, accuracy is known in many cases to poorly reflect the practical consequences of classification error, most famously in…

Statistics Theory · Mathematics 2022-09-27 Shashank Singh , Justin Khim

Macro-AUC is the arithmetic mean of the class-wise AUCs in multi-label learning and is commonly used in practice. However, its theoretical understanding is far lacking. Toward solving it, we characterize the generalization properties of…

Machine Learning · Computer Science 2023-06-05 Guoqiang Wu , Chongxuan Li , Yilong Yin

We address the problem of aggregating an ensemble of predictors with known loss bounds in a semi-supervised binary classification setting, to minimize prediction loss incurred on the unlabeled data. We find the minimax optimal predictions…

Machine Learning · Computer Science 2016-11-08 Akshay Balsubramani , Yoav Freund