Related papers: How Long Does Infinite Width Last? Signal Propagat…
The effectiveness of recurrent neural networks can be largely influenced by their ability to store into their dynamical memory information extracted from input sequences at different frequencies and timescales. Such a feature can be…
Networks of strongly-coupled neurons with random connectivity exhibit chaotic, asynchronous fluctuations. In previous work, we showed that when endowed with an additional low-rank connectivity consisting of the outer product of orthogonal…
We revisit the problem of extraordinary transmission of acoustic (electromagnetic) waves through a slit in a rigid (perfectly conducting) wall. We use matched asymptotic expansions to study the pertinent limit where the slit width is small…
Deeper modern architectures are costly to train, making hyperparameter transfer preferable to expensive repeated tuning. Maximal Update Parametrization ($\mu$P) helps explain why many hyperparameters transfer across width. Yet depth scaling…
After a more than decade-long period of relatively little research activity in the area of recurrent neural networks, several new developments will be reviewed here that have allowed substantial progress both in understanding and in…
We study how far a diffusion process on a graph can deviate from a designed starting pattern when the pattern is generated via Laplacian regularisation. Under standard stability conditions for undirected, entrywise nonnegative graphs, we…
Neural network width and depth are fundamental aspects of network topology. Universal approximation theorems provide that with increasing width or depth, there exists a neural network that approximates a function arbitrarily well. These…
We study the effects of mobility on two crucial characteristics in multi-scale dynamic networks: percolation and connection times. Our analysis provides insights into the question, to what extent long-time averages are well-approximated by…
The NTK is a widely used tool in the theoretical analysis of deep learning, allowing us to look at supervised deep neural networks through the lenses of kernel regression. Recently, several works have investigated kernel models for…
Recent work by Baratin et al. (2021) sheds light on an intriguing pattern that occurs during the training of deep neural networks: some layers align much more with data compared to other layers (where the alignment is defined as the…
We study free string propagation in families of plane wave geometries developing strong scale-invariant singularities in certain limits. We relate the singular limit of the evolution for all excited string modes to that of the…
Ability of deep networks to extract high level features and of recurrent networks to perform time-series inference have been studied. In view of universality of one hidden layer network at approximating functions under weak constraints, the…
We consider the problem of linear fitting of noisy data in the case of broad (say $\alpha$-stable) distributions of random impacts ("noise"), which can lack even the first moment. This situation, common in statistical physics of small…
A longstanding challenge for the Machine Learning community is the one of developing models that are capable of processing and learning from very long sequences of data. The outstanding results of Transformers-based networks (e.g., Large…
We numerically analyze the distribution of scattering resonance widths in one- and quasi-one dimensional tight binding models, in the localized regime. We detect and discuss an algebraic decay of the distribution, similar, though not…
Finite-width fully connected neural networks with Gaussian-initialized weights deviate from their infinite-width Gaussian limit, exhibiting non-vanishing higher-order cumulants. We approximate these deviations, for a neural network…
Using Stein's method techniques introduced by Chatterjee (2008) and further extended by Kasprzak and Peccati (2022) and by Lachi\`eze-Rey and Peccati (2017), we derive novel quantitative bounds on the convergence in distribution of…
Common to all different kinds of recurrent neural networks (RNNs) is the intention to model relations between data points through time. When there is no immediate relationship between subsequent data points (like when the data points are…
In this paper, we investigate how much of the numerical artefacts introduced by finite system size and choice of boundary conditions can be removed by finite size scaling, for strongly-correlated systems with quasi-long-range order.…
The evolution of a deep neural network trained by the gradient descent can be described by its neural tangent kernel (NTK) as introduced in [20], where it was proven that in the infinite width limit the NTK converges to an explicit limiting…