Related papers: Column randomization and almost-isometric embeddin…
We consider the empirical eigenvalue distribution for a class of non-Hermitian random block tridiagonal matrices $T$ with independent entries. The matrix has $n$ blocks on the diagonal and each block has size $\ell_n$, so the whole matrix…
A result of Simonovits and S\'os states that for any fixed graph $H$ and any $\epsilon > 0$ there exists $\delta > 0$ such that if $G$ is an $n$-vertex graph with the property that every $S \subseteq V(G)$ contains $p^{e(H)} |S|^{v(H)} \pm…
This paper investigates the spectral norm version of the column subset selection problem. Given a matrix $\mathbf{A}\in\mathbb{R}^{n\times d}$ and a positive integer $k\leq\text{rank}(\mathbf{A})$, the objective is to select exactly $k$…
We introduce kernel thinning, a new procedure for compressing a distribution $\mathbb{P}$ more effectively than i.i.d. sampling or standard thinning. Given a suitable reproducing kernel $\mathbf{k}_{\star}$ and $O(n^2)$ time, kernel…
Let $X_1,..., X_N\in\R^n$ be independent centered random vectors with log-concave distribution and with the identity as covariance matrix. We show that with overwhelming probability at least $1 - 3 \exp(-c\sqrt{n}\r)$ one has $ \sup_{x\in…
We construct a matrix $M\in R^{m\otimes d^c}$ with just $m=O(c\,\lambda\,\varepsilon^{-2}\text{poly}\log1/\varepsilon\delta)$ rows, which preserves the norm $\|Mx\|_2=(1\pm\varepsilon)\|x\|_2$ of all $x$ in any given $\lambda$ dimensional…
We consider approximation of diameter of a set $S$ of $n$ points in dimension $m$. E$\tilde{g}$ecio$\tilde{g}$lu and Kalantari \cite{kal} have shown that given any $p \in S$, by computing its farthest in $S$, say $q$, and in turn the…
The problem of finding large average submatrices of a real-valued matrix arises in the exploratory analysis of data from a variety of disciplines, ranging from genomics to social sciences. In this paper we provide a detailed asymptotic…
We investigate subsets with small sumset in arbitrary abelian groups. For an abelian group $G$ and an $n$-element subset $Y \subseteq G$ we show that if $m \ll s^2/(\log n)^2$, then the number of subsets $A \subseteq Y$ with $|A| = s$ and…
Data augmentation is one of the most popular techniques for improving the robustness of neural networks. In addition to directly training the model with original samples and augmented samples, a torrent of methods regularizing the distance…
Let $X=\Lambda\backslash\mathbb{H}$ be a Schottky surface, that is, a conformally compact hyperbolic surface of infinite area. Let $\delta$ denote the Hausdorff dimension of the limit set of $\Lambda$. We prove that for any compact subset…
We consider a convex constrained Gaussian sequence model and characterize necessary and sufficient conditions for the least squares estimator (LSE) to be minimax optimal. For a closed convex set $K\subset \mathbb{R}^n$ we observe…
Under certain conditions on k we calculate the limit distribution of the k:th largest eigenvalue, x_k, of the Gaussian Unitary Ensemble (GUE). More specifically, if n is the dimension of a random matrix from the GUE and k is such that both…
The L1-regularized Gaussian maximum likelihood estimator (MLE) has been shown to have strong statistical guarantees in recovering a sparse inverse covariance matrix, or alternatively the underlying graph structure of a Gaussian Markov…
We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min-Max Theorem (CGMT) to non-Gaussian settings, we derive an asymptotic min-max…
A family of random matrix ensembles interpolating between the GUE and the Ginibre ensemble of $n\times n$ matrices with iid centered complex Gaussian entries is considered. The asymptotic spectral distribution in these models is uniform in…
We give a general unified method that can be used for $L_1$ {\em closeness testing} of a wide range of univariate structured distribution families. More specifically, we design a sample optimal and computationally efficient algorithm for…
Let $N(L)$ be the number of eigenvalues, in an interval of length $L$, of a matrix chosen at random from the Gaussian Orthogonal, Unitary or Symplectic ensembles of ${\cal N}$ by ${\cal N}$ matrices, in the limit ${\cal…
We consider $N\times N$ Gaussian random matrices, whose average density of eigenvalues has the Wigner semi-circle form over $[-\sqrt{2},\sqrt{2}]$. For such matrices, using a Coulomb gas technique, we compute the large $N$ behavior of the…
We use the delta method and Stein's method to derive, under regularity conditions, explicit upper bounds for the distributional distance between the distribution of the maximum likelihood estimator (MLE) of a $d$-dimensional parameter and…