Related papers: How Long Does Infinite Width Last? Signal Propagat…
We study the statistics of level widths of a quantum dot with extended contacts in the absence of time-reversal symmetry. The widths are determined by the amplitude of the wavefunction averaged over the contact area. The distribution…
Modern deep networks are heavily overparameterized yet often generalize well, suggesting a form of low intrinsic complexity not reflected by parameter counts. We study this complexity at initialization through the effective rank of the…
Studying wave propagation in nonlinear discrete systems is essential for understanding energy transfer in lattices. While linear systems prohibit wave propagation within the natural band gap, nonlinear systems exhibit {supratransmission},…
The interplay between infinite-width neural networks (NNs) and classes of Gaussian processes (GPs) is well known since the seminal work of Neal (1996). While numerous theoretical refinements have been proposed in the recent years, the…
The dispersion of a diffusive scalar in a fluid flowing through a network has many applications including to biological flows, porous media, water supply and urban pollution. Motivated by this, we develop a large-deviation theory that…
We analyze the dynamics of finite width effects in wide but finite feature learning neural networks. Starting from a dynamical mean field theory description of infinite width deep neural network kernel and prediction dynamics, we provide a…
The way in which different types of dynamics unfold in complex networks is intrinsically related to the propagation of activation along nodes, which is strongly affected by the network connectivity. In this work we investigate to which…
Probabilistic graphical models are a powerful concept for modeling high-dimensional distributions. Besides modeling distributions, probabilistic graphical models also provide an elegant framework for performing statistical inference;…
This work theoretically studies stochastic neural networks, a main type of neural network in use. We prove that as the width of an optimized stochastic neural network tends to infinity, its predictive variance on the training set decreases…
Modern neural networks (NN) featuring a large number of layers (depth) and units per layer (width) have achieved a remarkable performance across many domains. While there exists a vast literature on the interplay between infinitely wide NNs…
In recent years, the mean field theory has been applied to the study of neural networks and has achieved a great deal of success. The theory has been applied to various neural network structures, including CNNs, RNNs, Residual networks, and…
We present a new unified theory of critical finite-size scaling for lattice statistical mechanical models with periodic boundary conditions above the upper critical dimension. Our theory is based on recent mathematically rigorous results…
Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon…
Given a network of fixed size $n$ and an initial distribution of data, we derive sufficient connectivity conditions on a sequence of time-varying digraphs for (a) data collection and (b) data dissemination, within at most $(n-1)$…
We analyze Elman-type Recurrent Reural Networks (RNNs) and their training in the mean-field regime. Specifically, we show convergence of gradient descent training dynamics of the RNN to the corresponding mean-field formulation in the large…
Machine learning models are increasingly trained or fine-tuned on synthetic data. Recursively training on such data has been observed to significantly degrade performance in a wide range of tasks, often characterized by a progressive drift…
The capability of recurrent neural networks to approximate trajectories of a random dynamical system, with random inputs, on non-compact domains, and over an indefinite or infinite time horizon is considered. The main result states that…
Guidance is a crucial technique for extracting the best performance out of image-generating diffusion models. Traditionally, a constant guidance weight has been applied throughout the sampling chain of an image. We show that guidance is…
Resonance trapping appears in open many-particle quantum systems at high level density when the coupling to the continuum of decay channels reaches a critical strength. Here a reorganization of the system takes place and a separation of…
Many stellar systems exhibit a finite spatial extent, yet constructing self-consistent spherical models with a prescribed outer boundary is non-trivial because sharp density cutoffs introduce discontinuities that lead to inconsistencies in…