Related papers: A Random Matrix Approach to Neural Networks
An $n \times n$ matrix with $\pm 1$ entries which acts on $\mathbb{R}^n$ as a scaled isometry is called Hadamard. Such matrices exist in some, but not all dimensions. Combining number-theoretic and probabilistic tools we construct matrices…
Recent works leveraging Graph Neural Networks to approach graph matching tasks have shown promising results. Recent progress in learning discrete distributions poses new opportunities for learning graph matching models. In this work, we…
The scaled standard Wigner matrix (symmetric with mean zero, variance one i.i.d. entries), and its limiting eigenvalue distribution, namely the semi-circular distribution, has attracted much attention. The $2k$th moment of the limit equals…
A central question in random matrix theory is universality. When an emergent phenomena is observed from a large collection of chosen random variables it is natural to ask if this behavior is specific to the chosen random variable or if the…
This paper examines the spectral characterizations of oriented graphs. Let $\Sigma$ be an $n$-vertex oriented graph with skew-adjacency matrix $S$. Previous research mainly focused on self-converse oriented graphs, proposing arithmetic…
Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on the restrictive assumption that the so-called feature…
Rank-width of a graph G, denoted by rw(G), is a width parameter of graphs introduced by Oum and Seymour (2006). We investigate the asymptotic behavior of rank-width of a random graph G(n,p). We show that, asymptotically almost surely, (i)…
Suppose $\alpha, \beta$ are Lipschitz strongly concave functions from $[0, 1]$ to $\mathbb{R}$ and $\gamma$ is a concave function from $[0, 1]$ to $\mathbb{R}$, such that $\alpha(0) = \gamma(0) = 0$, and $\alpha(1) = \beta(0) = 0$ and…
We study the first gradient descent step on the first-layer parameters $\boldsymbol{W}$ in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\top\sigma(\boldsymbol{W}^\top\boldsymbol{x})$, where…
Let $\mathcal{G}$ be a directed graph with vertices $1,2,\ldots, 2N$. Let $\mathcal{T}=(T_{i,j})_{(i,j)\in\mathcal{G}}$ be a family of contractive similitudes. For every $1\leq i\leq N$, let $i^+:=i+N$. For $1\leq i,j\leq N$, we define…
Graph Neural Networks (GNNs) have improved unsupervised community detection of clustered nodes due to their ability to encode the dual dimensionality of the connectivity and feature information spaces of graphs. Identifying the latent…
Randomized matrix sparsification has proven to be a fruitful technique for producing faster algorithms in applications ranging from graph partitioning to semidefinite programming. In the decade or so of research into this technique, the…
We introduce a general class of algorithms and supply a number of general results useful for analysing these algorithms when applied to regular graphs of large girth. As a result, we can transfer a number of results proved for random…
This paper studies the approximation and generalization abilities of score-based neural network generative models (SGMs) in estimating an unknown distribution $P_0$ from $n$ i.i.d. observations in $d$ dimensions. Assuming merely that $P_0$…
Thurstone's latent-normal model, introduced a century ago to describe human preferences in psychometrics (1927), remains a cornerstone for modeling random rankings. Yet when the underlying normals differ in distribution, the joint law of…
In this paper, we aim at recovering an unknown signal x0 from noisy L1measurements y=Phi*x0+w, where Phi is an ill-conditioned or singular linear operator and w accounts for some noise. To regularize such an ill-posed inverse problem, we…
We prove the universal asymptotically almost sure non-singularity of general Ginibre and Wigner ensembles of random matrices when the distribution of the entries are independent but not necessarily identically distributed and may depend on…
We derive upper bounds on the Wasserstein distance ($W_1$), with respect to $\sup$-norm, between any continuous $\mathbb{R}^d$ valued random field indexed by the $n$-sphere and the Gaussian, based on Stein's method. We develop a novel…
The purpose of this article is to study the eigenvalues $u_1^{\, t}=e^{it\theta_1},\dots,u_N^{\,t}=e^{it\theta_N}$ of $U^t$ where $U$ is a large $N\times N$ random unitary matrix and $t>0$. In particular we are interested in the typical…
Identification-robust hypothesis tests are commonly based on the continuous updating GMM objective function. When the number of moment conditions grows proportionally with the sample size, the large-dimensional weighting matrix prohibits…