English
Related papers

Related papers: Robust Mean Estimation With Auxiliary Samples

200 papers

This note examines the behavior of generalization capabilities - as defined by out-of-sample mean squared error (MSE) - of Linear Gaussian (with a fixed design matrix) and Linear Least Squares regression. Particularly, we consider a…

Statistics Theory · Mathematics 2021-09-21 Karthik Duraisamy

A lower bound on the minimum mean-squared error (MSE) in a Bayesian estimation problem is proposed in this paper. This bound utilizes a well-known connection to the deterministic estimation setting. Using the prior distribution, the bias…

Information Theory · Computer Science 2009-05-27 Zvika Ben-Haim , Yonina C. Eldar

The question of optimally approximating an arbitrary probability measure in the Wasserstein distance by a discrete one with uniform weights is considered. Estimates are obtained for the optimal approximation distance, with an explicit rate…

Probability · Mathematics 2026-04-14 Benjamin Seeger

This paper studies the expected optimal value of a mixed 0-1 programming problem with uncertain objective coefficients following a joint distribution. We assume that the true distribution is not known exactly, but a set of independent…

Optimization and Control · Mathematics 2017-08-28 Guanglin Xu , Samuel Burer

Towards understanding the fundamental limits of estimation from data of varied quality, we study the problem of estimating a mean parameter from heteroskedastic Gaussian observations where the variances are unknown and may vary arbitrarily…

Statistics Theory · Mathematics 2026-03-17 Yanjun Han , Abhishek Shetty , Jacob Shkrob

In the present study, we propose a new estimator for population mean of the study variable y in the case of stratified random sampling using the information based on auxiliary variable x. Expression for the mean squared error (MSE) of the…

Applications · Statistics 2014-07-25 Sachin Malik , Viplav Kumar Singh , Rajesh Singh

When a population exhibits heterogeneity, we often model it via a finite mixture: decompose it into several different but homogeneous subpopulations. Contemporary practice favors learning the mixtures by maximizing the likelihood for…

Machine Learning · Statistics 2021-07-06 Qiong Zhang , Jiahua Chen

We use distributionally-robust optimization for machine learning to mitigate the effect of data poisoning attacks. We provide performance guarantees for the trained model on the original data (not including the poison records) by training…

Machine Learning · Computer Science 2020-01-30 Farhad Farokhi

We propose a distributionally robust approach to risk-sensitive estimation of an unknown signal x from an observed signal y. The unknown signal and observation are modeled as random vectors whose joint probability distribution is unknown,…

Machine Learning · Computer Science 2026-04-21 Feras Al Taha , Eilyan Bitar

Self-supervised sequential recommendation significantly improves recommendation performance by maximizing mutual information with well-designed data augmentations. However, the mutual information estimation is based on the calculation of…

Machine Learning · Computer Science 2023-06-21 Ziwei Fan , Zhiwei Liu , Hao Peng , Philip S Yu

Minimum distance estimation (MDE) gained recent attention as a formulation of (implicit) generative modeling. It considers minimizing, over model parameters, a statistical distance between the empirical data distribution and the model. This…

Statistics Theory · Mathematics 2020-10-21 Ziv Goldfeld , Kristjan Greenewald , Kengo Kato

In this paper we tackle the problem of comparing distributions of random variables and defining a mean pattern between a sample of random events. Using barycenters of measures in the Wasserstein space, we propose an iterative version as an…

Statistics Theory · Mathematics 2013-12-12 Emmanuel Boissard , Thibaut Le Gouic , Jean-Michel Loubes

We refer to recent inference methodology and formulate a framework for solving the distributionally robust optimization problem, where the true probability measure is inside a Wasserstein ball around the empirical measure and the radius of…

Mathematical Finance · Quantitative Finance 2023-06-28 Xin Hai , Kihun Nam

Many randomized approximation algorithms operate by giving a procedure for simulating a random variable $X$ which has mean $\mu$ equal to the target answer, and a relative standard deviation bounded above by a known constant $c$. Examples…

Computation · Statistics 2019-08-16 Mark Huber

We introduce a distributionally robust maximum likelihood estimation model with a Wasserstein ambiguity set to infer the inverse covariance matrix of a $p$-dimensional Gaussian random vector from $n$ independent samples. The proposed model…

Optimization and Control · Mathematics 2018-05-21 Viet Anh Nguyen , Daniel Kuhn , Peyman Mohajerin Esfahani

The James-Stein estimator's dominance over maximum likelihood in terms of mean square error (MSE) has been one of the most celebrated results in modern statistics, suggesting that biased estimators can systematically outperform unbiased…

Statistics Theory · Mathematics 2025-08-12 Paul W. Vos

We analyze the complexity of sampling from a class of heavy-tailed distributions by discretizing a natural class of It\^o diffusions associated with weighted Poincar\'e inequalities. Based on a mean-square analysis, we establish the…

Statistics Theory · Mathematics 2023-03-03 Ye He , Tyler Farghly , Krishnakumar Balasubramanian , Murat A. Erdogdu

The adapted Wasserstein distance controls the calibration errors of optimal values in various stochastic optimization problems, pricing and hedging problems, optimal stopping problems, etc. However, statistical aspects of the adapted…

Probability · Mathematics 2025-09-16 Songyan Hou

For optimization on large-scale data, exactly calculating its solution may be computationally difficulty because of the large size of the data. In this paper we consider subsampled optimization for fast approximating the exact solution. In…

Machine Learning · Statistics 2018-04-11 Rong Zhu , Jiming Jiang

The Wasserstein metric is an important measure of distance between probability distributions, with applications in machine learning, statistics, probability theory, and data analysis. This paper provides upper and lower bounds on…

Statistics Theory · Mathematics 2019-11-11 Shashank Singh , Barnabás Póczos