English
Related papers

Related papers: Robust Max Statistics for High-Dimensional Inferen…

200 papers

Over the last two decades, many exciting variable selection methods have been developed for finding a small group of covariates that are associated with the response from a large pool. Can the discoveries from these data mining approaches…

Statistics Theory · Mathematics 2017-07-24 Jianqing Fan , Qi-Man Shao , Wen-Xin Zhou

Drawing statistical inferences from large datasets in a model-robust way is an important problem in statistics and data science. In this paper, we propose methods that are robust to large and unequal noise in different observational units…

Statistics Theory · Mathematics 2024-01-10 Edgar Dobriban , Weijie J. Su , Yachong Yang , Zhixiang Zhang

Bootstrap for nonlinear statistics like U-statistics of dependent data has been studied by several authors. This is typically done by producing a bootstrap version of the sample and plugging it into the statistic. We suggest an alternative…

Statistics Theory · Mathematics 2015-05-28 Olimjon Sh. Sharipov , Johannes Tewes , Martin Wendler

Concentration inequalities form an essential toolkit in the study of high dimensional (HD) statistical methods. Most of the relevant statistics literature in this regard is based on sub-Gaussian or sub-exponential tail assumptions. In this…

Statistics Theory · Mathematics 2023-01-09 Arun Kumar Kuchibhotla , Abhishek Chakrabortty

We consider the issue of performing accurate small sample inference in beta autoregressive moving average model, which is useful for modeling and forecasting continuous variables that assumes values in the interval $(0,1)$. The inferences…

Computation · Statistics 2017-02-16 Bruna Gregory Palm , Fábio M. Bayer

We study the problem of robustly estimating the mean of a $d$-dimensional distribution given $N$ examples, where most coordinates of every example may be missing and $\varepsilon N$ examples may be arbitrarily corrupted. Assuming each…

Data Structures and Algorithms · Computer Science 2021-05-04 Lunjia Hu , Omer Reingold

Simultaneous inference for high-dimensional non-Gaussian time series is always considered to be a challenging problem. Such tasks require not only robust estimation of the coefficients in the random process, but also deriving limiting…

Methodology · Statistics 2021-11-03 Linbo Liu , Danna Zhang

Analyzing high-dimensional data with manifold learning algorithms often requires searching for the nearest neighbors of all observations. This presents a computational bottleneck in statistical manifold learning when observations of…

Machine Learning · Computer Science 2022-03-11 Fan Cheng , Anastasios Panagiotelis , Rob J Hyndman

Data subject to heavy-tailed errors are commonly encountered in various scientific fields, especially in the modern era with explosion of massive data. To address this problem, procedures based on quantile regression and Least Absolute…

Statistics Theory · Mathematics 2014-10-09 Jianqing Fan , Quefeng Li , Yuyan Wang

We study the rate of convergence of the Mallows distance between the empirical distribution of a sample and the underlying population. The surprising feature of our results is that the convergence rate is slower in the discrete case than in…

Probability · Mathematics 2007-06-13 Richard Samworth , Oliver Johnson

We consider inference in linear regression models that is robust to heteroskedasticity and the presence of many control variables. When the number of control variables increases at the same rate as the sample size the usual…

Statistics Theory · Mathematics 2020-09-29 Koen Jochmans

Statistical methods for functional data are of interest for many applications. In this paper, we prove a central limit theorem for random variables taking their values in a Hilbert space. The random variables are assumed to be weakly…

Statistics Theory · Mathematics 2014-07-16 H. Dehling , O. Sh. Sharipov , M. Wendler

The validity of various bootstrapping methods has been proved for the sample mean of strongly mixing data. But in many applications, there appear nonlinear statistics of processes that are not strongly mixing. We investigate the…

Statistics Theory · Mathematics 2011-07-28 Olimjon Sh. Sharipov , Martin Wendler

This paper addresses the problem of estimating the tail index of distributions with heavy, Pareto-type tails for dependent data, that is of interest in the areas of finance, insurance, environmental monitoring and teletraffic analysis. A…

Methodology · Statistics 2010-05-25 Stilian A Stoev , George Michailidis

We benchmark the robustness of maximum likelihood based uncertainty estimation methods to outliers in training data for regression tasks. Outliers or noisy labels in training data results in degraded performances as well as incorrect…

Machine Learning · Computer Science 2022-02-09 Deebul S. Nair , Nico Hochgeschwender , Miguel A. Olivares-Mendez

In this paper we investigate how the bootstrap can be applied to time series regressions when the volatility of the innovations is random and non-stationary. The volatility of many economic and financial time series displays persistent…

Econometrics · Economics 2021-01-12 H. Peter Boswijk , Giuseppe Cavaliere , Anders Rahbek , Iliyan Georgiev

The existing theory of penalized quantile regression for longitudinal data has focused primarily on point estimation. In this work, we investigate statistical inference. We propose a wild residual bootstrap procedure and show that it is…

Econometrics · Economics 2022-05-10 Carlos Lamarche , Thomas Parker

This work is concerned with the limiting spectral distribution of rank-based dependency measures in high dimensions. We provide distribution-free results for multivariate empirical versions of Kendall's $\tau$ and Spearman's $\rho$ in a…

Statistics Theory · Mathematics 2025-08-22 Nina Dörnemann , Michael Fleermann , Johannes Heiny

In this paper we study a bootstrap strategy for estimating the variance of a mean taken over large multifactor crossed random effects data sets. We apply bootstrap reweighting independently to the levels of each factor, giving each…

Methodology · Statistics 2012-09-28 Art B. Owen , Dean Eckles

We commonly encounter the problem of identifying an optimally weight adjusted version of the empirical distribution of observed data, adhering to predefined constraints on the weights. Such constraints often manifest as restrictions on the…

Machine Learning · Statistics 2024-01-17 Abhisek Chakraborty , Anirban Bhattacharya , Debdeep Pati