English
Related papers

Related papers: Mean Testing under Truncation beyond Gaussian

200 papers

We develop conservative tests for the mean of a bounded population under stratified sampling and apply them to risk-limiting post-election audits. The tests are ``anytime valid'' under sequential sampling, allowing optional stopping in each…

Methodology · Statistics 2026-03-17 Jacob V. Spertus , Mayuri Sridhar , Philip B. Stark

We introduce a novel approach to finite sample robustness that avoids the pessimism of traditional breakdown analyses. We define the threshold breakdown point, the smallest contamination fraction needed to induce a prescribed deviation, and…

Statistics Theory · Mathematics 2026-05-19 Tianjun Ke , Marco Avella Medina

In this paper we consider the statistical inference of the unknown parameter of an exponential distribution based on the time truncated data. The time truncated data occurs quite often in the reliability analysis for type-I or hybrid…

Applications · Statistics 2017-03-06 Arnab Koley , Debasis Kundu

Suppose that we have $n$ agents and $n$ items which lie in a shared metric space. We would like to match the agents to items such that the total distance from agents to their matched items is as small as possible. However, instead of having…

Computer Science and Game Theory · Computer Science 2023-05-23 Nima Anari , Moses Charikar , Prasanna Ramakrishnan

We provide optimal lower bounds for two well-known parameter estimation (also known as statistical estimation) tasks in high dimensions with approximate differential privacy. First, we prove that for any $\alpha \le O(1)$, estimating the…

Statistics Theory · Mathematics 2024-01-05 Shyam Narayanan

Given a randomized experiment with binary outcomes, exact confidence intervals for the average causal effect of the treatment can be computed through a series of permutation tests. This approach requires minimal assumptions and is valid for…

Methodology · Statistics 2025-06-19 P. M. Aronow , Haoge Chang , Patrick Lopatto

Learning high-dimensional distributions is a significant challenge in machine learning and statistics. Classical research has mostly concentrated on asymptotic analysis of such data under suitable assumptions. While existing works…

Machine Learning · Computer Science 2024-11-19 Sutanu Gayen , Sanket Kale , Sayantan Sen

A common problem in genetics is that of testing whether a set of highly dependent gene expressions differ between two populations, typically in a high-dimensional setting where the data dimension is larger than the sample size. Most…

Methodology · Statistics 2015-03-11 Måns Thulin

We study high-dimensional asymptotic performance limits of binary supervised classification problems where the class conditional densities are Gaussian with unknown means and covariances and the number of signal dimensions scales faster…

Machine Learning · Statistics 2016-11-17 Mohammad Hossein Rohban , Prakash Ishwar , Birant Orten , William C. Karl , Venkatesh Saligrama

Under certain conditions, the largest eigenvalue of a sample covariance matrix undergoes a well-known phase transition when the sample size $n$ and data dimension $p$ diverge proportionally. In the subcritical regime, this eigenvalue has…

Statistics Theory · Mathematics 2025-04-01 Nina Dörnemann , Miles E. Lopes

The sample mean is often used to aggregate different unbiased estimates of a parameter, producing a final estimate that is unbiased but possibly high-variance. This paper introduces the Bayesian median of means, an aggregation rule that…

Statistics Theory · Mathematics 2019-06-05 Paulo Orenstein

We study a class of hypothesis testing problems in which, upon observing the realization of an $n$-dimensional Gaussian vector, one has to decide whether the vector was drawn from a standard normal distribution or, alternatively, whether…

Statistics Theory · Mathematics 2010-11-22 Louigi Addario-Berry , Nicolas Broutin , Luc Devroye , Gábor Lugosi

Heteroscedasticity -- where the variance of a variable changes with other variables -- is pervasive in real data, and elucidating why it arises from the perspective of statistical moments is crucial in scientific knowledge discovery and…

Machine Learning · Statistics 2026-05-28 Yoichi Chikahara

This paper studies the truncation method from Alquier [1] to derive high-probability PAC-Bayes bounds for unbounded losses with heavy tails. Assuming that the $p$-th moment is bounded, the resulting bounds interpolate between a slow rate $1…

Machine Learning · Statistics 2024-03-26 Borja Rodríguez-Gálvez , Omar Rivasplata , Ragnar Thobaben , Mikael Skoglund

In contrast to the empirical mean, the Median-of-Means (MoM) is an estimator of the mean $\theta$ of a square integrable r.v. $Z$, around which accurate nonasymptotic confidence bounds can be built, even when $Z$ does not exhibit a…

Machine Learning · Statistics 2021-02-09 Pierre Laforgue , Guillaume Staerman , Stephan Clémençon

We investigate the sample complexity of mutual information and conditional mutual information testing. For conditional mutual information testing, given access to independent samples of a triple of random variables $(A, B, C)$ with unknown…

Data Structures and Algorithms · Computer Science 2025-06-05 Jan Seyfried , Sayantan Sen , Marco Tomamichel

In this paper, we consider the magnetic anomaly detection problem which aims to find hidden ferromagnetic masses by estimating the weak perturbation they induce on local Earth's magnetic field. We consider classical detection schemes that…

Signal Processing · Electrical Eng. & Systems 2025-11-10 Clément Chenevas-Paule , Steeve Zozor , Laure-Line Rouve , Olivier J. J. Michel , Olivier Pinaud , Romain Kukla

We use bias-reduced estimators of high quantiles, of heavy-tailed distributions, to introduce a new estimator of the mean in the case of infinite second moment. The asymptotic normality of the proposed estimator is established and checked,…

Methodology · Statistics 2014-05-09 Brahim Brahimi , Djamel Meraghni , Abdelhakim Necir , Djabrane Yahia

Traditionally, robust statistics has focused on designing estimators tolerant to a minority of contaminated data. Robust list-decodable learning focuses on the more challenging regime where only a minority $\frac 1 k$ fraction of the…

Data Structures and Algorithms · Computer Science 2020-11-20 Ilias Diakonikolas , Daniel M. Kane , Daniel Kongsgaard , Jerry Li , Kevin Tian

Nonparametric two-sample tests such as the Maximum Mean Discrepancy (MMD) are often used to detect differences between two distributions in machine learning applications. However, the majority of existing literature assumes that error-free…

Machine Learning · Statistics 2023-08-08 Ron Nafshi , Maggie Makar