English
Related papers

Related papers: Finite-sample analysis of M-estimators using self-…

200 papers

A key feature of a sequential study is that the actual sample size is a random variable that typically depends on the outcomes collected. While hypothesis testing theory for sequential designs is well established, parameter and precision…

Statistics Theory · Mathematics 2017-12-21 Ben Berckmoes , Geert Molenberghs

Motivated by real-world machine learning applications, we analyze approximations to the non-asymptotic fundamental limits of statistical classification. In the binary version of this problem, given two training sequences generated according…

Information Theory · Computer Science 2018-12-07 Lin Zhou , Vincent Y. F. Tan , Mehul Motani

For classification problems with significant class imbalance, subsampling can reduce computational costs at the price of inflated variance in estimating model parameters. We propose a method for subsampling efficiently for logistic…

Computation · Statistics 2014-09-24 William Fithian , Trevor Hastie

We seek an entropy estimator for discrete distributions with fully empirical accuracy bounds. As stated, this goal is infeasible without some prior assumptions on the distribution. We discover that a certain information moment assumption…

Information Theory · Computer Science 2022-12-27 Doron Cohen , Aryeh Kontorovich , Aaron Koolyk , Geoffrey Wolfer

Importance sampling is a popular variance reduction method for Monte Carlo estimation, where a notorious question is how to design good proposal distributions. While in most cases optimal (zero-variance) estimators are theoretically…

Statistics Theory · Mathematics 2021-02-22 Carsten Hartmann , Lorenz Richter

In observational studies, accurately characterizing variance is critical for sample size determination, yet unaccounted-for variability from propensity score estimation and the resulting weights limit the accuracy of standard variance…

Methodology · Statistics 2026-04-24 Taekwon Hong , Daeyoung Lim , Woojung Bae , Yong Ma

Let $\mathcal{G}$ be a directed graph with vertices $1,2,\ldots, 2N$. Let $\mathcal{T}=(T_{i,j})_{(i,j)\in\mathcal{G}}$ be a family of contractive similitudes. For every $1\leq i\leq N$, let $i^+:=i+N$. For $1\leq i,j\leq N$, we define…

Functional Analysis · Mathematics 2023-07-12 Sanguo Zhu

In constrained stochastic optimization, one naturally expects that imposing a stricter feasible set does not increase the statistical risk of an estimator defined by projection onto that set. In this paper, we show that this intuition can…

Statistics Theory · Mathematics 2026-01-23 Omar Al-Ghattas

This paper presents uniform estimation and inference theory for a large class of nonparametric partitioning-based M-estimators. The main theoretical results include: (i) uniform consistency for convex and non-convex objective functions;…

Statistics Theory · Mathematics 2025-09-01 Matias D. Cattaneo , Yingjie Feng , Boris Shigida

The likelihood ratio statistic, with its asymptotic $\chi^2$ distribution at regular model points, is often used for hypothesis testing. At model singularities and boundaries, however, the asymptotic distribution may not be $\chi^2$, as…

Statistics Theory · Mathematics 2018-06-25 Jonathan D. Mitchell , Elizabeth S. Allman , John A. Rhodes

When the experimental data set is contaminated, we usually employ robust alternatives to common location and scale estimators such as the sample median and Hodges-Lehmann estimators for location and the sample median absolute deviation and…

Methodology · Statistics 2020-08-11 Chanseok Park , Haewon Kim , Min Wang

Overparameterized models fail to generalize well in the presence of data imbalance even when combined with traditional techniques for mitigating imbalances. This paper focuses on imbalanced classification datasets, in which a small subset…

Machine Learning · Computer Science 2022-06-28 Tina Behnia , Ke Wang , Christos Thrampoulidis

We propose a general method for constructing confidence intervals and statistical tests for single or low-dimensional components of a large parameter vector in a high-dimensional model. It can be easily adjusted for multiplicity taking…

Statistics Theory · Mathematics 2014-06-24 Sara van de Geer , Peter Bühlmann , Ya'acov Ritov , Ruben Dezeure

In this paper, we develop a non-asymptotic local normal approximation for multinomial probabilities. First, we use it to find non-asymptotic total variation bounds between the measures induced by uniformly jittered multinomials and the…

Statistics Theory · Mathematics 2023-09-06 Eric Bax , Frédéric Ouimet

We study prediction and estimation problems using empirical risk minimization, relative to a general convex loss function. We obtain sharp error rates even when concentration is false or is very restricted, for example, in heavy-tailed…

Machine Learning · Statistics 2014-10-14 Shahar Mendelson

It is well known that, under standard regularity conditions, the maximum likelihood estimator (MLE) satisfies a central limit theorem and converges in distribution to a Gaussian random variable as the sample size grows. This paper…

Information Theory · Computer Science 2026-05-26 Leighton P. Barnes , Alex Dytso

We propose a new definition of the chi-square divergence between distributions. Based on convexity properties and duality, this version of the {\chi}^2 is well suited both for the classical applications of the {\chi}^2 for the analysis of…

Statistics Theory · Mathematics 2011-01-26 Michel Broniatowski , Samantha Leorato

Existing theories on deep nonparametric regression have shown that when the input data lie on a low-dimensional manifold, deep neural networks can adapt to the intrinsic data structures. In real world applications, such an assumption of…

Machine Learning · Computer Science 2023-06-27 Zixuan Zhang , Minshuo Chen , Mengdi Wang , Wenjing Liao , Tuo Zhao

We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…

Machine Learning · Computer Science 2024-12-03 Maryam Aliakbarpour , Piotr Indyk , Ronitt Rubinfeld , Sandeep Silwal

In Ruckdeschel[10], we derive an asymptotic expansion of the maximal mean squared error (MSE) of location M-estimators on suitably thinned out, shrinking gross error neighborhoods. In this paper, we compile several consequences of this…

Statistics Theory · Mathematics 2010-06-02 Peter Ruckdeschel