Related papers: Uniform Convergence Beyond Glivenko-Cantelli
Heisenberg's uncertainty principle was originally posed for the limit of the accuracy of simultaneous measurement of non-commuting observables as stating that canonically conjugate observables can be measured simultaneously only with the…
This paper addresses the statistical problem of estimating the infinite-norm deviation from the empirical mean to the distribution mean for high-dimensional distributions on $\{0,1\}^d$, potentially with $d=\infty$. Unlike traditional…
This article considers the parametric estimation of $Pr(X<Y<Z)$ and its generalizations based on several well-known one-parameter and two-parameter continuous distributions. It is shown that for some one-parameter distributions and when…
Creating accurate meta-embeddings from pre-trained source embeddings has received attention lately. Methods based on global and locally-linear transformation and concatenation have shown to produce accurate meta-embeddings. In this paper,…
It is common to assume in empirical research that observables and unobservables are additively separable, especially, when the former are endogenous. This is done because it is widely recognized that identification and estimation challenges…
Ensembles are a straightforward, remarkably effective method for improving the accuracy,calibration, and robustness of models on classification tasks; yet, the reasons that underlie their success remain an active area of research. We build…
For many applications, an ensemble of base classifiers is an effective solution. The tuning of its parameters(number of classes, amount of data on which each classifier is to be trained on, etc.) requires G, the generalization error of a…
We establish a central limit theorem for the unnormalized linear statistic of the Gaussian Unitary Ensemble under optimal conditions: the linear statistics converges if and only if the expression for the limiting variance is finite.
Consider informative selection of a sample from a finite population. Responses are realized as independent and identically distributed (i.i.d.) random variables with a probability density function (p.d.f.) f, referred to as the…
Randomness is ubiquitous in many applications across data science and machine learning. Remarkably, systems composed of random components often display emergent global behaviors that appear deterministic, manifesting a transition from…
Statistical learning theory chiefly studies restricted hypothesis classes, particularly those with finite Vapnik-Chervonenkis (VC) dimension. The fundamental quantity of interest is the sample complexity: the number of samples required to…
In this paper, we develop a computational approach for estimating the mean value of a quantity in the presence of uncertainty. We demonstrate that, under some mild assumptions, the upper and lower bounds of the mean value are efficiently…
We describe various moment-based ensemble interpretation models for the construction of probabilistic temperature forecasts from ensembles. We apply the methods to one year of medium range ensemble forecasts and perform in and out of sample…
New Vapnik and Chervonenkis type concentration inequalities are derived for the empirical distribution of an independent random sample. Focus is on the maximal deviation over classes of Borel sets within a low probability region. The…
Solomonoff's uncomputable universal prediction scheme $\xi$ allows to predict the next symbol $x_k$ of a sequence $x_1...x_{k-1}$ for any Turing computable, but otherwise unknown, probabilistic environment $\mu$. This scheme will be…
We investigate a generalized empirical likelihood approach in a two-group setting where the constraints on parameters have a form of U-statistics. In this situation, the summands that consist of the constraints for the empirical likelihood…
The consistency of a learning method is usually established under the assumption that the observations are a realization of an independent and identically distributed (i.i.d.) or mixing process. Yet, kernel methods such as support vector…
We extend the study of \emph{melonic} quartic tensor models to models with arbitrary quartic interactions. This extension requires a new version of the loop vertex expansion using several species of intermediate fields and iterated…
We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…
Recent advances in center-based clustering continue to improve upon the drawbacks of Lloyd's celebrated $k$-means algorithm over $60$ years after its introduction. Various methods seek to address poor local minima, sensitivity to outliers,…