English
Related papers

Related papers: On the Concentration of the Missing Mass

200 papers

In machine learning models, the estimation of errors is often complex due to distribution bias, particularly in spatial data such as those found in environmental studies. We introduce an approach based on the ideas of importance sampling to…

Machine Learning · Computer Science 2023-09-15 Boris Prokhorov , Diana Koldasbayeva , Alexey Zaytsev

Let $(X,Y)\in\mathcal{X}\times \mathcal{Y}$ be a random couple with unknown distribution $P$. Let $\GG$ be a class of measurable functions and $\ell$ a loss function. The problem of statistical learning deals with the estimation of the…

Statistics Theory · Mathematics 2012-07-12 Sébastien Loustau

We study full Bayesian procedures for sparse linear regression when errors have a symmetric but otherwise unknown distribution. The unknown error distribution is endowed with a symmetrized Dirichlet process mixture of Gaussians. For the…

Statistics Theory · Mathematics 2019-03-26 Minwoo Chae , Lizhen Lin , David B. Dunson

We consider the problem of recovering a distribution function on the real line from observations additively contaminated with errors following the standard Laplace distribution. Assuming that the latent distribution is completely unknown…

Methodology · Statistics 2017-08-21 Catia Scricciolo

We consider the problem of estimating rare event probabilities, focusing on systems whose evolution is governed by differential equations with uncertain input parameters. If the system dynamics is expensive to compute, standard sampling…

Computation · Statistics 2019-11-05 Siddhant Wahal , George Biros

We investigate Bayesian predictive inference for finite population quantities when there are unequal probabilities of selection. Only limited information about the sample design is available; i.e., only the first-order selection…

Methodology · Statistics 2018-04-10 Junheng Ma , Joe Sedransk , Balgobin Nandram , Lu Chen

We provide a systematic approach to deal with the following problem. Let $X_1,\ldots,X_n$ be, possibly dependent, $[0,1]$-valued random variables. What is a sharp upper bound on the probability that their sum is significantly larger than…

Probability · Mathematics 2015-07-27 Christos Pelekis , Jan Ramon

We study a missing-value imputation method, termed kNNSampler, that imputes a given unit's missing response by randomly sampling from the observed responses of the $k$ most similar units to the given unit in terms of the observed…

Machine Learning · Statistics 2025-12-03 Parastoo Pashmchi , Jérôme Benoit , Motonobu Kanagawa

We consider a collection of independent random variables that are identically distributed, except for a small subset which follows a different, anomalous distribution. We study the problem of detecting which random variables in the…

Information Theory · Computer Science 2018-06-21 Natalie Durgin , Rachel Grotheer , Chenxi Huang , Shuang Li , Anna Ma , Deanna Needell , Jing Qin

The probability of the small deviations of the matrix $AA^T$ determinant is estimated, where $A$ is an $n\times\infty$ random matrix with centered entries having joint Gaussian distribution. The inequality obtained is sharp in a sence.

Probability · Mathematics 2013-03-19 Nadezhda V. Volodko

The Ewens sampling formula is a distribution related to the random partition of a positive integer. In this study, we investigate the issue of non-existence solutions in parameter estimation under the distribution. As a result, the first…

Statistics Theory · Mathematics 2021-05-25 Masayo Y. Hirose , Shuhei Mano

We studied two probabilistic models of the distribution of primes in the natural number [1].The paper considers the third probabilistic model of the distribution of primes in the natural number. The author proved that the results obtained…

Number Theory · Mathematics 2015-09-30 Victor Volfson

Several proofs of the monotonicity of the non-Gaussianness (divergence with respect to a Gaussian random variable with identical second order statistics) of the sum of n independent and identically distributed (i.i.d.) random variables were…

Information Theory · Computer Science 2007-07-13 Jacob Binia

In the context of a species sampling problem we discuss a non-parametric maximum likelihood estimator for the underlying probability mass function. The estimator is known in the computer science literature as the high profile estimator. We…

Statistics Theory · Mathematics 2018-01-12 Dragi Anevski , Richard D. Gill , Stefan Zohren

We propose a methodology for modeling and comparing probability distributions within a Bayesian nonparametric framework. Building on dependent normalized random measures, we consider a prior distribution for a collection of discrete random…

Methodology · Statistics 2022-06-01 Mario Beraha , Jim E. Griffin

Dynamic factor models are often estimated by point-estimation methods, disregarding parameter uncertainty. We propose a method accounting for parameter uncertainty by means of posterior approximation, using variational inference. Our…

Methodology · Statistics 2022-10-14 Erik Spånberg

Necessary and sufficient conditions of uniform consistency are explored. A hypothesis is simple. Nonparametric sets of alternatives are bounded convex sets in $\mathbb{L}_p$, $p >1$ with "small" balls deleted. The "small" balls have the…

Statistics Theory · Mathematics 2024-03-07 Mikhail Ermakov

Missing data can lead to inefficiencies and biases in analyses, in particular when data are missing not at random (MNAR). It is thus vital to understand and correctly identify the missing data mechanism. Recovering missing values through a…

Methodology · Statistics 2022-12-08 Jack Noonan , Adetola Adedamola Adediran , Robin Mitra , Stefanie Biedermann

Let $\a$ be a real-valued random variable of mean zero and variance 1. Let $M_n(\a)$ denote the $n \times n$ random matrix whose entries are iid copies of $\a$ and $\sigma_n(M_n(\a))$ denote the least singular value of $M_n(\a)$.…

Probability · Mathematics 2009-03-04 Terence Tao , Van Vu

Missing Not At Random (MNAR) values lead to significant biases in the data, since the probability of missingness depends on the unobserved values.They are ''not ignorable'' in the sense that they often require defining a model for the…

Statistics Theory · Mathematics 2020-06-11 Aude Sportisse , Claire Boyer , Julie Josse