English
Related papers

Related papers: Eigenvalues of Autoencoders in Training and at Ini…

200 papers

We identify incremental learning dynamics in transformers, where the difference between trained and initial weights progressively increases in rank. We rigorously prove this occurs under the simplifying assumptions of diagonal weight…

Machine Learning · Computer Science 2023-12-12 Enric Boix-Adsera , Etai Littwin , Emmanuel Abbe , Samy Bengio , Joshua Susskind

In this paper, we investigate the eigenvalue distribution of a class of kernel random matrices whose $(i,j)$-th entry is $f(X_i,X_j)$ where $f$ is a symmetric function belonging to the Paley-Wiener space $\mathcal{B}_c$ and $(X_i)_{1\leq i…

Statistics Theory · Mathematics 2025-07-22 Jebalia Mohamed , Ahmed Souabni

Unsupervised pre-training was a critical technique for training deep neural networks years ago. With sufficient labeled data and modern training techniques, it is possible to train very deep neural networks from scratch in a purely…

Computer Vision and Pattern Recognition · Computer Science 2017-03-29 Jianfeng Dong , Xiao-Jiao Mao , Chunhua Shen , Yu-Bin Yang

Consider a $N\times n$ random matrix $Z_n=(Z^n_{j_1 j_2})$ where the individual entries are a realization of a properly rescaled stationary gaussian random field. The purpose of this article is to study the limiting empirical distribution…

Probability · Mathematics 2007-06-13 W. Hachem , P. Loubaton , J. Najim

Good initialization is essential for training Deep Neural Networks (DNNs). Oftentimes such initialization is found through a trial and error approach, which has to be applied anew every time an architecture is substantially modified, or…

Machine Learning · Statistics 2022-06-29 Tianyu He , Darshil Doshi , Andrey Gromov

Convolutional autoencoders are now at the forefront of image compression research. To improve their entropy coding, encoder output is typically analyzed with a second autoencoder to generate per-variable parametrized prior probability…

Image and Video Processing · Electrical Eng. & Systems 2021-11-18 Benoit Brummer , Christophe De Vleeschouwer

Current approaches to learning vector representations of text that are compatible between different languages usually require some amount of parallel text, aligned at word, sentence or at least document level. We hypothesize however, that…

Computation and Language · Computer Science 2016-08-11 Antonio Valerio Miceli Barone

Does a Variational AutoEncoder (VAE) consistently encode typical samples generated from its decoder? This paper shows that the perhaps surprising answer to this question is `No'; a (nominally trained) VAE does not necessarily amortize…

Machine Learning · Computer Science 2020-12-08 A. Taylan Cemgil , Sumedh Ghaisas , Krishnamurthy Dvijotham , Sven Gowal , Pushmeet Kohli

We present in this paper a novel approach for training deterministic auto-encoders. We show that by adding a well chosen penalty term to the classical reconstruction cost function, we can achieve results that equal or surpass those attained…

Artificial Intelligence · Computer Science 2011-04-22 Salah Rifai , Xavier Muller , Xavier Glorot , Gregoire Mesnil , Yoshua Bengio , Pascal Vincent

What do auto-encoders learn about the underlying data generating distribution? Recent work suggests that some auto-encoder variants do a good job of capturing the local manifold structure of data. This paper clarifies some of these previous…

Machine Learning · Computer Science 2014-08-20 Guillaume Alain , Yoshua Bengio

We derive the distribution of the eigenvalues of a large sample covariance matrix when the data is dependent in time. More precisely, the dependence for each variable $i=1,...,p$ is modelled as a linear process…

Probability · Mathematics 2012-01-19 Oliver Pfaffel , Eckhard Schlemm

Convolutional Neural Networks spread through computer vision like a wildfire, impacting almost all visual tasks imaginable. Despite this, few researchers dare to train their models from scratch. Most work builds on one of a handful of…

Computer Vision and Pattern Recognition · Computer Science 2016-09-26 Philipp Krähenbühl , Carl Doersch , Jeff Donahue , Trevor Darrell

There are individual differences in expressive behaviors driven by cultural norms and personality. This between-person variation can result in reduced emotion recognition performance. Therefore, personalization is an important step in…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Minh Tran , Yufeng Yin , Mohammad Soleymani

A new algorithmic framework is proposed for learning autoencoders of data distributions. We minimize the discrepancy between the model and target distributions, with a \emph{relational regularization} on the learnable latent prior. This…

Machine Learning · Computer Science 2020-06-29 Hongteng Xu , Dixin Luo , Ricardo Henao , Svati Shah , Lawrence Carin

It has become common practice now to use random initialization schemes, rather than the pre-trained embeddings, when training transformer based models from scratch. Indeed, we find that pre-trained word embeddings from GloVe, and some…

Computation and Language · Computer Science 2024-07-18 Ha Young Kim , Niranjan Balasubramanian , Byungkon Kang

We look at the eigenvalues of the Hessian of a loss function before and after training. The eigenvalue distribution is seen to be composed of two parts, the bulk which is concentrated around zero, and the edges which are scattered away from…

Machine Learning · Computer Science 2017-10-06 Levent Sagun , Leon Bottou , Yann LeCun

We analyze the training of a two-layer autoencoder used to parameterize a flow-based generative model for sampling from a high-dimensional Gaussian mixture. Previous work shows that the phase where the relative probability between the modes…

Machine Learning · Computer Science 2025-02-11 Santiago Aranguri , Francesco Insulla

We consider $N\times N$ Hermitian or symmetric random matrices with independent entries. The distribution of the $(i,j)$-th matrix element is given by a probability measure $\nu_{ij}$ whose first two moments coincide with those of the…

Mathematical Physics · Physics 2011-11-16 Antti Knowles , Jun Yin

The Koopman autoencoder, a data-driven technique, has gained traction for modeling nonlinear dynamics using deep learning methods in recent years. Given the linear characteristics inherent to the Koopman operator, controlling its…

Machine Learning · Computer Science 2024-08-22 Jinho Choi , Sivaram Krishnan , Jihong Park

In this article, we will look at autoencoders. This article covers the mathematics and the fundamental concepts of autoencoders. We will discuss what they are, what the limitations are, the typical use cases, and we will look at some…

Machine Learning · Computer Science 2022-01-12 Umberto Michelucci