Related papers: Eigenvalues of Autoencoders in Training and at Ini…
Two-term asymptotic formulae for the probability distribution functions for the smallest eigenvalue of the Jacobi $ \beta $-Ensembles are derived for matrices of large size in the r\'egime where $ \beta > 0 $ is arbitrary and one of the…
Deep neural networks are capable of modelling highly non-linear functions by capturing different levels of abstraction of data hierarchically. While training deep networks, first the system is initialized near a good optimum by greedy…
The ability of Variational Autoencoders (VAEs) to learn disentangled representations has made them popular for practical applications. However, their behaviour is not yet fully understood. For example, the questions of when they can provide…
Semantic segmentation labels are expensive and time consuming to acquire. Hence, pretraining is commonly used to improve the label-efficiency of segmentation models. Typically, the encoder of a segmentation model is pretrained as a…
Feedforward neural networks are widely used as universal predictive models to fit data distribution. Common gradient-based learning, however, suffers from many drawbacks making the training process ineffective and time-consuming.…
We consider some random band matrices with band-width $N^\mu$ whose entries are independent random variables with distribution tail in $x^{-\alpha}$. We consider the largest eigenvalues and the associated eigenvectors and prove the…
Recent studies have shown that many important aspects of neural network learning take place within the very earliest iterations or epochs of training. For example, sparse, trainable sub-networks emerge (Frankle et al., 2019), gradient…
In the present work, eigenvalue distributions defined by a random rectangular matrix whose components are neither independently nor identically distributed are analyzed using replica analysis and belief propagation. In particular, we…
In order to comprehend and enhance models that describes various brain regions is important to study the dynamics of trained recurrent neural networks. Including Dales law in such models usually presents several challenges. However, this is…
Autoencoders exhibit impressive abilities to embed the data manifold into a low-dimensional latent space, making them a staple of representation learning methods. However, without explicit supervision, which is often unavailable, the…
Different brain areas, such as the cortex and, more specifically, the prefrontal cortex, show great recurrence in their connections, even in early sensory areas. {Several approaches and methods based on trained networks have been proposed…
We construct custom regularization functions for use in supervised training of deep neural networks. Our technique is applicable when the ground-truth labels themselves exhibit internal structure; we derive a regularizer by learning an…
The distribution of the ratios of consecutive eigenvalue spacings of random matrices has emerged as an important tool to study spectral properties of many-body systems. This article numerically investigates the eigenvalue ratios…
This paper centers on the limit eigenvalue distribution for random Vandermonde matrices with unit magnitude complex entries. The phases of the entries are chosen independently and identically distributed from the interval $[-\pi,\pi]$.…
We analyze gene co-expression network under the random matrix theory framework. The nearest neighbor spacing distribution of the adjacency matrix of this network follows Gaussian orthogonal statistics of random matrix theory (RMT). Spectral…
The point of this paper is to question typical assumptions in deep learning and suggest alternatives. A particular contribution is to prove that even if a Stacked Convolutional Auto-Encoder is good at reconstructing pictures, it is not…
The point of this paper is to question typical assumptions in deep learning and suggest alternatives. A particular contribution is to prove that even if a Stacked Convolutional Auto-Encoder is good at reconstructing pictures, it is not…
In recent years, transformer-based models have revolutionized deep learning, particularly in sequence modeling. To better understand this phenomenon, there is a growing interest in using Markov input processes to study transformers.…
We study the convergence properties of a pair of learning algorithms (learning with and without memory). This leads us to study the dominant eigenvalue of a class of random matrices. This turns out to be related to the roots of the…
We study random matrices acting on tensor product spaces which have been transformed by a linear block operation. Using operator-valued free probability theory, under some mild assumptions on the linear map acting on the blocks, we compute…