English
Related papers

Related papers: How Long Does Infinite Width Last? Signal Propagat…

200 papers

A longstanding goal in deep learning research has been to precisely characterize training and generalization. However, the often complex loss landscapes of neural networks have made a theory of learning dynamics elusive. In this work, we…

We study the gradient-based training of large-depth residual networks (ResNets) from standard random initializations. We show that infinite-depth ResNets behave as if they were infinitely wide, regardless of their actual width. More…

Machine Learning · Computer Science 2026-03-04 Lénaïc Chizat

Recent developments in applications of artificial neural networks with over $n=10^{14}$ parameters make it extremely important to study the large $n$ behaviour of such networks. Most works studying wide neural networks have focused on the…

Machine Learning · Computer Science 2023-04-10 Luís Carvalho , João Lopes Costa , José Mourão , Gonçalo Oliveira

We study the effect of width on the dynamics of feature-learning neural networks across a variety of architectures and datasets. Early in training, wide neural networks trained on online data have not only identical loss curves but also…

Machine Learning · Computer Science 2023-12-07 Nikhil Vyas , Alexander Atanasov , Blake Bordelon , Depen Morwani , Sabarish Sainathan , Cengiz Pehlevan

To better understand the temporal characteristics and the lifetime of fluctuations in stochastic processes in networks, we investigated diffusive persistence in various graphs. Global diffusive persistence is defined as the fraction of…

Statistical Mechanics · Physics 2024-06-04 Omar Malik , Melinda Varga , Alaa Moussawi , David Hunt , Boleslaw Szymanski , Zoltan Toroczkai , Gyorgy Korniss

In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime'…

Machine Learning · Computer Science 2025-05-09 Qiyang Han , Masaaki Imaizumi

One of the most influential results in neural network theory is the universal approximation theorem [1, 2, 3] which states that continuous functions can be approximated to within arbitrary accuracy by single-hidden-layer feedforward neural…

Machine Learning · Computer Science 2021-12-16 Clemens Hutter , Recep Gül , Helmut Bölcskei

There is a growing literature on the study of large-width properties of deep Gaussian neural networks (NNs), i.e. deep NNs with Gaussian-distributed parameters or weights, and Gaussian stochastic processes. Motivated by some empirical and…

Machine Learning · Computer Science 2023-04-11 Alberto Bordino , Stefano Favaro , Sandra Fortini

Light propagation in optical waveguides with periodically modulated index of refraction and alternating gain and loss are investigated for linear and nonlinear systems. Based on a multiscale perturbation analysis, it is shown that for many…

Optics · Physics 2015-06-23 Sean Nixon , Jianke Yang

Wide neural networks have proven to be a rich class of architectures for both theory and practice. Motivated by the observation that finite width convolutional networks appear to outperform infinite width networks, we study scaling laws for…

Machine Learning · Computer Science 2020-08-21 Anders Andreassen , Ethan Dyer

We study the behavior of untrained neural networks whose weights and biases are randomly distributed using mean field theory. We show the existence of depth scales that naturally limit the maximum depth of signal propagation through these…

Machine Learning · Statistics 2017-04-06 Samuel S. Schoenholz , Justin Gilmer , Surya Ganguli , Jascha Sohl-Dickstein

We study the distributional properties of linear neural networks with random parameters in the context of large networks, where the number of layers diverges in proportion to the number of neurons per layer. Prior works have shown that in…

Machine Learning · Statistics 2024-11-26 Federico Bassetti , Lucia Ladelli , Pietro Rotondo

Vanishing (and exploding) gradients effect is a common problem for recurrent neural networks with nonlinear activation functions which use backpropagation method for calculation of derivatives. Deep feedforward neural networks with many…

Neural and Evolutionary Computing · Computer Science 2017-02-15 Artem Chernodub , Dimitri Nowicki

Nonlinear stripe patterns occur in many different systems, from the small scales of biological cells to geological scales as cloud patterns. They all share the universal property of being stable at different wavenumbers $q$, i.e., they are…

Pattern Formation and Solitons · Physics 2022-02-22 Mirko Ruppert , Walter Zimmermann

We develop a mathematically rigorous framework for multilayer neural networks in the mean field regime. As the network's widths increase, the network's learning trajectory is shown to be well captured by a meaningful and dynamically…

Machine Learning · Computer Science 2023-02-14 Phan-Minh Nguyen , Huy Tuan Pham

In this paper, we explore the persistent current in thin superconducting wires and accurately examine the effects of the phase slips on that current. The main result of the paper is the formula for persistent current in terms of the…

Mesoscale and Nanoscale Physics · Physics 2018-07-10 Ilya Vilkoviskiy

In this paper, we derive theoretical bounds for the long-term influence of a node in an Independent Cascade Model (ICM). We relate these bounds to the spectral radius of a particular matrix and show that the behavior is sub-critical when…

Probability · Mathematics 2014-07-18 Remi Lemonnier , Kevin Scaman , Nicolas Vayatis

Unstable homoepitaxy on rough substrates is treated within a linear continuum theory. The time dependence of the surface width W(t) is governed by three length scales: The characteristic scale $l_0$ of the substrate roughness, the terrace…

Statistical Mechanics · Physics 2009-10-31 Joachim Krug , Martin Rost

Stability is a fundamental property of dynamical systems, yet to this date it has had little bearing on the practice of recurrent neural networks. In this work, we conduct a thorough investigation of stable recurrent models. Theoretically,…

Machine Learning · Computer Science 2019-03-05 John Miller , Moritz Hardt

Comparing Bayesian neural networks (BNNs) with different widths is challenging because, as the width increases, multiple model properties change simultaneously, and, inference in the finite-width case is intractable. In this work, we…

Machine Learning · Statistics 2022-11-29 Jiayu Yao , Yaniv Yacoby , Beau Coker , Weiwei Pan , Finale Doshi-Velez