English
Related papers

Related papers: Jackknife Variance Estimation for H\'ajek-Dominate…

200 papers

We show that a simple modification of the 1-nearest neighbor classifier yields a strongly Bayes consistent learner. Prior to this work, the only strongly Bayes consistent proximity-based method was the k-nearest neighbor classifier, for k…

Machine Learning · Computer Science 2018-08-20 Aryeh Kontorovich , Roi Weiss

We provide algorithms for regression with adversarial responses under large classes of non-i.i.d. instance sequences, on general separable metric spaces, with provably minimal assumptions. We also give characterizations of learnability in…

Machine Learning · Computer Science 2023-06-13 Moïse Blanchard , Patrick Jaillet

Motivated by small bandwidth asymptotics for kernel-based semiparametric estimators in econometrics, this paper establishes Gaussian approximation results for high-dimensional fixed-order $U$-statistics whose kernels depend on the sample…

Statistics Theory · Mathematics 2025-10-15 Shunsuke Imai , Yuta Koike

Random forests remain among the most popular off-the-shelf supervised learning algorithms. Despite their well-documented empirical success, however, until recently, few theoretical results were available to describe their performance and…

Machine Learning · Statistics 2021-11-17 Wei Peng , Tim Coleman , Lucas Mentch

In this article we discuss estimation of the common variance of several normal populations with tree order restricted means. We discuss the asymptotic properties of the maximum likelihood estimator of the variance as the number of…

Statistics Theory · Mathematics 2014-07-24 Antar Bandyopadhyay , Sanjay Chaudhuri

Modern statistical analysis often encounters datasets with large sizes. For these datasets, conventional estimation methods can hardly be used immediately because practitioners often suffer from limited computational resources. In most…

Methodology · Statistics 2023-04-14 Shuyuan Wu , Xuening Zhu , Hansheng Wang

Given a random sample from a multivariate population, estimating the number of large eigenvalues of the population covariance matrix is an important problem in Statistics with wide applications in many areas. In the context of Principal…

Statistics Theory · Mathematics 2020-11-10 Abhinav Chakraborty , Soumendu Sundar Mukherjee , Arijit Chakrabarti

This paper investigates weighted approximations for studentized $U$-statistics type processes, both with symmetric and antisymmetric kernels, only under the assumption that the distribution of the projection variate is in the domain of…

Probability · Mathematics 2007-11-12 Miklós Csörgő , Barbara Szyszkowicz , Qiying Wang

The classical theory of rank-based inference is entirely based either on ordinary ranks, which do not allow for considering location (intercept) parameters, or on signed ranks, which require an assumption of symmetry. If the median, in the…

Statistics Theory · Mathematics 2007-06-13 Marc Hallin , Catherine Vermandele , Bas Werker

Heavy-tailed distributions, such as the Cauchy distribution, are acknowledged for providing more accurate models for financial returns, as the normal distribution is deemed insufficient for capturing the significant fluctuations observed in…

Statistics Theory · Mathematics 2025-07-31 Ganesh Vishnu Avhad , Ananya Lahiri , Sudheesh K. Kattumannil

The error or variability of machine learning algorithms is often assessed by repeatedly re-fitting a model with different weighted versions of the observed data. The ubiquitous tools of cross-validation (CV) and the bootstrap are examples…

Methodology · Statistics 2020-02-10 Ryan Giordano , Will Stephenson , Runjing Liu , Michael I. Jordan , Tamara Broderick

We propose a Hausman test for the correct specification of unobserved heterogeneity in both linear and nonlinear fixed-effects panel data models. The null hypothesis is that heterogeneity is either time-invariant or, symmetrically,…

Econometrics · Economics 2025-09-03 Claudia Pigini , Alessandro Pionati , Francesco Valentini

A Lorenz curve is a graphical representation of the distribution of income or wealth within a population. The generalized Lorenz curve can be created by scaling the values on the vertical axis of a Lorenz curve by the average output of the…

Methodology · Statistics 2023-09-26 Suthakaran Ratnasingam , Anton Butenko

Network experiments are powerful tools for studying spillover effects, which avoid endogeneity by randomly assigning treatments to units over networks. However, it is non-trivial to analyze network experiments properly without imposing…

Econometrics · Economics 2025-06-09 Mengsi Gao , Peng Ding

Balanced repeated replication (BRR) and the jackknife are two widely used methods for estimating variances in stratified samples with two primary sampling units per stratum. While both methods produce variance estimators that can be…

Methodology · Statistics 2026-03-13 Matthias von Davier

This paper studies the Gaussian and bootstrap approximations for the probabilities of a non-degenerate U-statistic belonging to the hyperrectangles in $\mathbb{R}^d$ when the dimension $d$ is large. A two-step Gaussian approximation…

Statistics Theory · Mathematics 2017-07-11 Xiaohui Chen

We propose new model selection criteria based on generalized ridge estimators dominating the maximum likelihood estimator under the squared risk and the Kullback-Leibler risk in multivariate linear regression. Our model selection criteria…

Statistics Theory · Mathematics 2016-04-08 Yuichi Mori , Taiji Suzuki

We derive high-probability finite-sample uniform rates of consistency for $k$-NN regression that are optimal up to logarithmic factors under mild assumptions. We moreover show that $k$-NN regression adapts to an unknown lower intrinsic…

Machine Learning · Statistics 2018-11-06 Heinrich Jiang

In this paper, we study the asymptotic bias of the factor-augmented regression estimator and its reduction, which is augmented by the $r$ factors extracted from a large number of $N$ variables with $T$ observations. In particular, we…

Methodology · Statistics 2025-10-02 Peiyun Jiang , Yoshimasa Uematsu , Takashi Yamagata

This paper addresses the following question: given a sample of i.i.d. random variables with finite variance, can one construct an estimator of the unknown mean that performs nearly as well as if the data were normally distributed? One of…

Statistics Theory · Mathematics 2023-02-06 Stanislav Minsker