Related papers: Robust covariance estimation under $L_4-L_2$ norm …
A common pursuit in modern statistical learning is to attain satisfactory generalization out of the source data distribution (OOD). In theory, the challenge remains unsolved even under the canonical setting of covariate shift for the linear…
Covariance matrix estimation is one of the most important problems in statistics. To accommodate the complexity of modern datasets, it is desired to have estimation procedures that not only can incorporate the structural assumptions of…
This note describes the concentration phenomenon for a high dimensional sub-gaussian vector \( X \). In the Gaussian case, for any linear operator \( Q \), it holds \( P\bigl( \| Q X \|^{2} - tr (B) > 2 \sqrt{x\, tr(B^{2})} + 2 \| B \| x…
Approximating significance scans of searches for new particles in high-energy physics experiments as Gaussian fields is a well-established way to estimate the trials factors required to quantify global significances. We propose a novel,…
We study the problem of high-dimensional sparse mean estimation in the presence of an $\epsilon$-fraction of adversarial outliers. Prior work obtained sample and computationally efficient algorithms for this task for identity-covariance…
Stacked lensing is a powerful means of measuring the average mass distribution around large-scale structure tracers. There are two stacked lensing estimators used in the literature, denoted as $\Delta\Sigma$ and $\gamma_+$, which are…
This paper studies the Gaussian approximation of high-dimensional and non-degenerate U-statistics of order two under the supremum norm. We propose a two-step Gaussian approximation procedure that does not impose structural assumptions on…
Gradient-based solvers risk convergence to local optima, leading to incorrect researcher inference. Heuristic-based algorithms are able to ``break free" of these local optima to eventually converge to the true global optimum. However, given…
We investigate robust linear regression where data may be contaminated by an oblivious adversary, i.e., an adversary than may know the data distribution but is otherwise oblivious to the realizations of the data samples. This model has been…
Most linear experimental design problems assume homogeneous variance although heteroskedastic noise is present in many realistic settings. Let a learner have access to a finite set of measurement vectors $\mathcal{X}\subset \mathbb{R}^d$…
We obtain robust and computationally efficient estimators for learning several linear models that achieve statistically optimal convergence rate under minimal distributional assumptions. Concretely, we assume our data is drawn from a…
We revisit heavy-tailed corrupted least-squares linear regression assuming to have a corrupted $n$-sized label-feature sample of at most $\epsilon n$ arbitrary outliers. We wish to estimate a $p$-dimensional parameter $b^*$ given such…
In this paper, we consider the problem of estimating the $p\times p$ scale matrix $\Sigma$ of a multivariate linear regression model $Y=X\,\beta + \mathcal{E}\,$ when the distribution of the observed matrix $Y$ belongs to a large class of…
Robust estimation is much more challenging in high dimensions than it is in one dimension: Most techniques either lead to intractable optimization problems or estimators that can tolerate only a tiny fraction of errors. Recent work in…
For normal canonical models with $X \sim N_p(\theta, \sigma^{2} I_{p}), \;\; S^{2} \sim \sigma^{2}\chi^{2}_{k}, \;{independent}$, we consider the problem of estimating $\theta$ under scale invariant squared error loss $\frac{\|d-\theta…
We consider the problem of finding an approximate solution to $\ell_1$ regression while only observing a small number of labels. Given an $n \times d$ unlabeled data matrix $X$, we must choose a small set of $m \ll n$ rows to observe the…
We consider the extreme eigenvalues of the sample covariance matrix $Q=YY^*$ under the generalized elliptical model that $Y=\Sigma^{1/2}XD.$ Here $\Sigma$ is a bounded $p \times p$ positive definite deterministic matrix representing the…
A multivariable measurement error model $AX \approx B$ is considered. Here $A$ and $B$ are input and output matrices of measurements and $X$ is a rectangular matrix of fixed size to be estimated. The errors in $[A,B]$ are row-wise…
In a recent paper the author obtained optimal bounds for the strong Gaussian approximation of sums of independent $\R^d$-valued random vectors with finite exponential moments. The results may be considered as generalizations of well-known…
Factor models are a very efficient way to describe high dimensional vectors of data in terms of a small number of common relevant factors. This problem, which is of fundamental importance in many disciplines, is usually reformulated in…