English
Related papers

Related papers: A fast and slightly robust covariance estimator

200 papers

Compositional data arise in many areas of research in the natural and biomedical sciences. One prominent example is in the study of the human gut microbiome, where one can measure the relative abundance of many distinct microorganisms in a…

Methodology · Statistics 2024-04-26 Aaron J. Molstad , Karl Oskar Ekvall , Piotr M. Suder

Estimation of the covariance matrix has attracted a lot of attention of the statistical research community over the years, partially due to important applications such as Principal Component Analysis. However, frequently used empirical…

Statistics Theory · Mathematics 2018-06-19 Stanislav Minsker

We introduce PseudoNet, a new pseudolikelihood-based estimator of the inverse covariance matrix, that has a number of useful statistical and computational properties. We show, through detailed experiments with synthetic and also real-world…

Methodology · Statistics 2016-10-17 Alnur Ali , Kshitij Khare , Sang-Yun Oh , Bala Rajaratnam

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample…

Methodology · Statistics 2015-11-24 Rong Zhu , Ping Ma , Michael W. Mahoney , Bin Yu

Given a non-negative random variable $W$ and $\theta>0$, let the generalized Dickman transformation map the distribution of $W$ to that of $$ W^*=_d U^{1/\theta}(W+1), $$ where $U \sim {\cal U}[0,1]$, a uniformly distributed variable on the…

Probability · Mathematics 2018-10-22 Larry Goldstein

This paper generalizes a part of the theory of $Z$-estimation which has been developed mainly in the context of modern empirical processes to the case of stochastic processes, typically, semimartingales. We present a general theorem to…

Statistics Theory · Mathematics 2009-09-03 Yoichi Nishiyama

Given i.i.d. observations of a random vector $X \in \mathbb{R}^p$, we study the problem of estimating both its covariance matrix $\Sigma^*$, and its inverse covariance or concentration matrix {$\Theta^* = (\Sigma^*)^{-1}$.} We estimate…

Machine Learning · Statistics 2008-11-24 Pradeep Ravikumar , Martin J. Wainwright , Garvesh Raskutti , Bin Yu

We consider the problem of finding an approximate solution to $\ell_1$ regression while only observing a small number of labels. Given an $n \times d$ unlabeled data matrix $X$, we must choose a small set of $m \ll n$ rows to observe the…

Machine Learning · Computer Science 2021-05-21 Aditya Parulekar , Advait Parulekar , Eric Price

Statistical estimation and inference for marginal hazard models with varying coefficients for multivariate failure time data are important subjects in survival analysis. A local pseudo-partial likelihood procedure is proposed for estimating…

Statistics Theory · Mathematics 2009-09-29 Jianwen Cai , Jianqing Fan , Haibo Zhou , Yong Zhou

We present a simple perturbation mechanism for the release of $d$-dimensional covariance matrices $\Sigma$ under pure differential privacy. For large datasets with at least $n\geq d^2/\varepsilon$ elements, our mechanism recovers the…

Machine Learning · Computer Science 2026-02-03 Tommaso d'Orsi , Gleb Novikov

Subsampling methods have been recently proposed to speed up least squares estimation in large scale settings. However, these algorithms are typically not robust to outliers or corruptions in the observed covariates. The concept of influence…

Machine Learning · Statistics 2014-06-20 Brian McWilliams , Gabriel Krummenacher , Mario Lucic , Joachim M. Buhmann

Let $p$ be an unknown and arbitrary probability distribution over $[0,1)$. We consider the problem of {\em density estimation}, in which a learning algorithm is given i.i.d. draws from $p$ and must (with high probability) output a…

Machine Learning · Computer Science 2014-11-04 Siu-On Chan , Ilias Diakonikolas , Rocco A. Servedio , Xiaorui Sun

Assume that $(X_t)_{t\in\Z}$ is a real valued time series admitting a common marginal density $f$ with respect to Lebesgue's measure. Donoho {\it et al.} (1996) propose a near-minimax method based on thresholding wavelets to estimate $f$ on…

Statistics Theory · Mathematics 2011-03-17 Irène Gannaz , Olivier Wintenberger

This work provides a unified analysis of the properties of the sample covariance matrix $\Sigma_n$ over the class of $p\times p$ population covariance matrices $\Sigma$ of reduced effective rank $r_e(\Sigma)$. This class includes scaled…

Statistics Theory · Mathematics 2015-06-02 Florentina Bunea , Luo Xiao

We provide a numerical scheme to approximate as closely as desired the Gaussian or exponential measure $\mu(\om)$ of (not necessarily compact) basic semi-algebraic sets$\om\subset\R^n$. We obtain two monotone (non increasing and non…

Optimization and Control · Mathematics 2017-07-11 Jean-Bernard Lasserre

We derive an upper bound for the efficiency of estimating entries in the inverse covariance matrix of a high dimensional distribution. We show that in order to approximate an off-diagonal entry of the density matrix of a $d$-dimensional…

Statistics Theory · Mathematics 2015-05-06 Ronen Eldan

We present multivariate unbiased estimators for second, third, and fourth order cumulants $C_2(x,y)$, $C_3(x,y,z)$, and $C_4(x,y,z,w)$. Many relevant new estimators are derived for cases where some variables are average-free or pairs of…

Statistics Theory · Mathematics 2019-04-30 Fabian Schefczik , Daniel Hägele

This paper proposes a new robust smooth-threshold estimating equation to select important variables and automatically estimate parameters for high dimensional longitudinal data. A novel working correlation matrix is proposed to capture…

Methodology · Statistics 2021-11-30 Liya Fu , Jiaqi Li , You-Gan Wang

We study the problem of PAC learning $\gamma$-margin halfspaces in the presence of Massart noise. Without computational considerations, the sample complexity of this learning problem is known to be $\widetilde{\Theta}(1/(\gamma^2…

Machine Learning · Computer Science 2025-01-17 Ilias Diakonikolas , Nikos Zarifis

We explore why many recently proposed robust estimation problems are efficiently solvable, even though the underlying optimization problems are non-convex. We study the loss landscape of these robust estimation problems, and identify the…

Machine Learning · Statistics 2020-05-29 Banghua Zhu , Jiantao Jiao , Jacob Steinhardt