English
Related papers

Related papers: Sample Complexity Bounds for Robust Mean Estimatio…

200 papers

As the use of machine learning in high impact domains becomes widespread, the importance of evaluating safety has increased. An important aspect of this is evaluating how robust a model is to changes in setting or population, which…

Machine Learning · Computer Science 2021-03-16 Adarsh Subbaswamy , Roy Adams , Suchi Saria

Uncertainty quantification for image data is dominated by complex deep learning methods, yet the field lacks an interpretable, mathematically grounded baseline. We propose Bayesian scattering to fill this gap, serving as a first-step…

Machine Learning · Computer Science 2026-03-24 Bernardo Fichera , Zarko Ivkovic , Kjell Jorner , Philipp Hennig , Viacheslav Borovitskiy

We propose an estimator for the mean of random variables in separable real Banach spaces using the empirical characteristic function. Assuming that the covariance operator of the random variable is bounded in a precise sense, we show that…

Statistics Theory · Mathematics 2020-11-04 Sohail Bahmani

We consider the problem of mean estimation under quantization and adversarial corruption. We construct multivariate robust estimators that are optimal up to logarithmic factors in two different settings. The first is a one-bit setting,…

Machine Learning · Statistics 2026-01-13 Pedro Abdalla , Junren Chen

Nonparametric estimation of a mixing distribution based on data coming from a mixture model is a challenging problem. Beyond estimation, there is interest in uncertainty quantification, e.g., confidence intervals for features of the mixing…

Methodology · Statistics 2019-06-14 Vaidehi Dixit , Ryan Martin

Normal mixture distributions are arguably the most important mixture models, and also the most technically challenging. The likelihood function of the normal mixture model is unbounded based on a set of random samples, unless an artificial…

Statistics Theory · Mathematics 2009-08-25 Jiahua Chen , Pengfei Li

The quantification problem consists of determining the prevalence of a given label in a target population. However, one often has access to the labels in a sample from the training population but not in the target population. A common…

Machine Learning · Statistics 2019-04-08 Afonso Fernandes Vaz , Rafael Izbicki , Rafael Bassi Stern

The estimation of information measures of continuous distributions based on samples is a fundamental problem in statistics and machine learning. In this paper, we analyze estimates of differential entropy in $K$-dimensional Euclidean space,…

Information Theory · Computer Science 2021-11-29 Georg Pichler , Pablo Piantanida , Günther Koliander

We study the problem of robust estimation under heterogeneous corruption rates, where each sample may be independently corrupted with a known but non-identical probability. This setting arises naturally in distributed and federated…

Machine Learning · Computer Science 2025-10-02 Syomantak Chaudhuri , Jerry Li , Thomas A. Courtade

A distributed inference scheme which uses bounded transmission functions over a Gaussian multiple access channel is considered. When the sensor measurements are decreasingly reliable as a function of the sensor index, the conditions on the…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-06-16 Sivaraman Dasarathan , Cihan Tepedelenlioglu

Specimens are collected from $N$ different sources. Each specimen has probability $p$ of being contaminated (e.g., in the case of an infectious disease, $p$ is the prevalence rate), independently of the other specimens. In many cases group…

Probability · Mathematics 2024-01-31 Vassilis G. Papanicolaou

As predictive algorithms grow in popularity, using the same dataset to both train and test a new model has become routine across research, policy, and industry. Sample-splitting attains valid inference on model properties by using separate…

Econometrics · Economics 2025-11-27 Bruno Fava

Group-invariant probability distributions appear in many data-generative models in machine learning, such as graphs, point clouds, and images. In practice, one often needs to estimate divergences between such distributions. In this work, we…

Machine Learning · Computer Science 2026-02-05 Behrooz Tahmasebi , Stefanie Jegelka

For a sample of absolutely bounded i.i.d. random variables with a continuous density the cumulative distribution function of the sample variance is represented by a univariate integral over a Fourier series. If the density is a polynomial…

Statistics Theory · Mathematics 2008-10-10 T. Royen

In this work we tackle the problem of estimating the density $f_X$ of a random variable $X$ by successive smoothing, such that the smoothed random variable $Y$ fulfills $(\partial_t - \Delta_1)f_Y(\,\cdot\,, t) = 0$, $f_Y(\,\cdot\,, 0) =…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Martin Zach , Thomas Pock , Erich Kobler , Antonin Chambolle

Coarse data arise when learners observe only partial information about samples; namely, a set containing the sample rather than its exact value. This occurs naturally through measurement rounding, sensor limitations, and lag in economic…

Machine Learning · Computer Science 2026-02-27 Alkis Kalavasis , Anay Mehrotra , Manolis Zampetakis , Felix Zhou , Ziyu Zhu

It is well known that if the power spectral density of a continuous time stationary stochastic process does not have a compact support, data sampled from that process at any uniform sampling rate leads to biased and inconsistent spectrum…

Statistics Theory · Mathematics 2010-06-09 Radhendushka Srivastava , Debasis Sengupta

Due to the complexity of order statistics, the finite sample behaviour of robust statistics is generally not analytically solvable. While the Monte Carlo method can provide approximate solutions, its convergence rate is typically very slow,…

Methodology · Statistics 2024-09-12 Li Tuobang

We introduce a rigorous and sensitive significance test for hyperuniformity that yields reliable results even from a single sample. Our approach is based on a detailed analysis of the empirical Fourier transform of a stationary point…

Statistics Theory · Mathematics 2026-03-23 Michael A. Klatt , Günter Last , Norbert Henze

With the advent of surveys containing millions to billions of galaxies, it is imperative to develop analysis techniques that utilize the available statistical power. In galaxy clustering, even small sample contamination arising from…

Cosmology and Nongalactic Astrophysics · Physics 2020-02-18 Humna Awan , Eric Gawiser