Related papers: Stochastic weight matrix dynamics during learning …
We consider a variant of the stochastic gradient descent (SGD) with a random learning rate and reveal its convergence properties. SGD is a widely used stochastic optimization algorithm in machine learning, especially deep learning. Numerous…
We employ constraints to control the parameter space of deep neural networks throughout training. The use of customized, appropriately designed constraints can reduce the vanishing/exploding gradients problem, improve smoothness of…
For a large class of feature maps we provide a tight asymptotic characterisation of the test error associated with learning the readout layer, in the high-dimensional limit where the input dimension, hidden layer widths, and number of…
A considerable number of systems have recently been reported in which Brownian yet non-Gaussian dynamics was observed. These are processes characterised by a linear growth in time of the mean squared displacement, yet the probability…
The Restricted Boltzmann Machine (RBM), an important tool used in machine learning in particular for unsupervized learning tasks, is investigated from the perspective of its spectral properties. Starting from empirical observations, we…
We analyze gradient descent with randomly weighted data points in a linear regression model, under a generic weighting distribution. This includes various forms of stochastic gradient descent, importance sampling, but also extends to…
Deep learning methods relying on multi-layered networks have been actively studied in a wide range of fields in recent years, and deep Boltzmann machines(DBMs) is one of them. In this study, a model of DBMs with some properites of weight…
We study the effect of mini-batching on the loss landscape of deep neural networks using spiked, field-dependent random matrix theory. We demonstrate that the magnitude of the extremal values of the batch Hessian are larger than those of…
Eugene Wigner's revolutionary vision predicted that the energy levels of large complex quantum systems exhibit a universal behavior: the statistics of energy gaps depend only on the basic symmetry type of the model. Simplified models of…
Stochastic models of diffusion with excluded-volume effects are used to model many biological and physical systems at a discrete level. The average properties of the population may be described by a continuum model based on partial…
We describe the application of tools from statistical mechanics to analyse the dynamics of various classes of supervised learning rules in perceptrons. The character of this paper is mostly that of a cross between a biased non-encyclopedic…
Recently, we introduced the active Dyson Brownian motion model (DBM), in which $N$ run-and-tumble particles interact via a logarithmic repulsive potential in the presence of a harmonic well. We found that in a broad range of parameters the…
The stochastic rotational invariance of an integration by parts formula inspired by the Bismut approach to Malliavin calculus is proved in the framework of the Lie symmetry theory of stochastic differential equations. The non-trivial effect…
We present a modified Brownian motion model for random matrices where the eigenvalues (or levels) of a random matrix evolve in "time" in such a way that they never cross each other's path. Also, owing to the exact integrability of the level…
Since the introduction of Dyson's Brownian motion in early 1960's, there have been a lot of developments in the investigation of stochastic processes on the space of Hermitian matrices. Their properties, especially, the properties of their…
For optimizing a non-convex function in finite dimension, a method is to add Brownian noise to a gradient descent, allowing for transitions between basins of attractions of different minimizers. To adapt this for optimization over a space…
Analyzing neural network dynamics via stochastic gradient descent (SGD) is crucial to building theoretical foundations for deep learning. Previous work has analyzed structured inputs within the \textit{hidden manifold model}, often under…
As an extension of the theory of Dyson's Brownian motion models for the standard Gaussian random-matrix ensembles, we report a systematic study of hermitian matrix-valued processes and their eigenvalue processes associated with the chiral…
We analyze the landscape and training dynamics of diagonal linear networks in a linear regression task, with the network parameters being perturbed by small isotropic normal noise. The addition of such noise may be interpreted as a…
Stochastic gradient optimization is the dominant learning paradigm for a variety of scenarios, from classical supervised learning to modern self-supervised learning. We consider stochastic gradient algorithms for learning problems whose…