Related papers: Minimum sample size for detection of Gutenberg-Ric…
Knowledge of the noise distribution in magnitude diffusion MRI images is the centerpiece to quantify uncertainties arising from the acquisition process. The use of parallel imaging methods, the number of receiver coils and imaging filters…
The minimizers sampling mechanism is a popular mechanism for string sampling introduced independently by Schleimer et al. [SIGMOD 2003] and by Roberts et al. [Bioinf. 2004]. Given two positive integers $w$ and $k$, it selects the…
Minimum divergence estimators provide a natural choice of estimators in a statistical inference problem. Different properties of various families of these divergence measures such as Hellinger distance, power divergence, density power…
Concentration inequalities for the sample mean, like those due to Bernstein, Hoeffding, and Bentkus, are valid for any sample size but overly conservative, yielding confidence intervals that are unnecessarily wide. The central limit theorem…
We introduce a new method for performing clustering with the aim of fitting clusters with different scatters and weights. It is designed by allowing to handle a proportion $\alpha$ of contaminating data to guarantee the robustness of the…
Recent advances in generative modeling have led to an increased interest in the study of statistical divergences as means of model comparison. Commonly used evaluation methods, such as the Frechet Inception Distance (FID), correlate well…
We revisit the problem of estimating the mean of a real-valued distribution, presenting a novel estimator with sub-Gaussian convergence: intuitively, "our estimator, on any distribution, is as accurate as the sample mean is for the Gaussian…
We describe and analyze a variance reduction approach for Monte Carlo (MC) sampling that accelerates the estimation of statistics of computationally expensive simulation models using an ensemble of models with lower cost. These lower cost…
We introduce a new method to measure the dispersion of mmax values of star clusters and show that the observed sample of mmax is inconsistent with random sampling from an universal stellar initial mass function (IMF) at a 99.9% confidence…
Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…
The positive false discovery rate (pFDR) is a useful overall measure of errors for multiple hypothesis testing, especially when the underlying goal is to attain one or more discoveries. Control of pFDR critically depends on how much…
Downsampling or under-sampling is a technique that is utilized in the context of large and highly imbalanced classification models. We study optimal downsampling for imbalanced classification using generalized linear models (GLMs). We…
Few shot learning is an important problem in machine learning as large labelled datasets take considerable time and effort to assemble. Most few-shot learning algorithms suffer from one of two limitations- they either require the design of…
We investigate how well the redshift distribution of a population of extragalactic objects can be reconstructed using angular cross-correlations with a sample whose redshifts are known. We derive the minimum variance quadratic estimator,…
Estimating prevalence, the fraction of a population with a certain medical condition, is fundamental to epidemiology. Traditional methods rely on classification of test samples taken at random from a population. Such approaches to…
The distribution of seismic moment is of capital interest to evaluate earthquake hazard, in particular regarding the most extreme events. We make use of likelihood-ratio tests to compare the simple Gutenberg-Richter power-law distribution…
Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data,…
We compute quantitative bounds for measuring the discrepancy between the distribution of two min-max statistics involving either pairs of Gaussian random matrices, or one Gaussian and one Gaussian-subordinated random matrix. In the fully…
In this paper, we study the problem of learning one-dimensional Gaussian mixture models (GMMs) with a specific focus on estimating both the model order and the mixing distribution from independent and identically distributed (i.i.d.)…
We study the discrepancy between the distribution of a vector-valued functional of i.i.d. random elements and that of a Gaussian vector. Our main contribution is an explicit bound on the convex distance between the two distributions,…