中文
相关论文

相关论文: How Long Does Infinite Width Last? Signal Propagat…

200 篇论文

The effectiveness of recurrent neural networks can be largely influenced by their ability to store into their dynamical memory information extracted from input sequences at different frequencies and timescales. Such a feature can be…

机器学习 · 计算机科学 2020-07-01 Antonio Carta , Alessandro Sperduti , Davide Bacciu

Networks of strongly-coupled neurons with random connectivity exhibit chaotic, asynchronous fluctuations. In previous work, we showed that when endowed with an additional low-rank connectivity consisting of the outer product of orthogonal…

神经元与认知 · 定量生物学 2021-06-09 Itamar Daniel Landau , Haim Sompolinsky

We revisit the problem of extraordinary transmission of acoustic (electromagnetic) waves through a slit in a rigid (perfectly conducting) wall. We use matched asymptotic expansions to study the pertinent limit where the slit width is small…

光学 · 物理学 2018-12-19 Jacob R. Holley , Ory Schnitzer

Deeper modern architectures are costly to train, making hyperparameter transfer preferable to expensive repeated tuning. Maximal Update Parametrization ($\mu$P) helps explain why many hyperparameters transfer across width. Yet depth scaling…

机器学习 · 计算机科学 2026-02-10 Shenxi Wu , Haosong Zhang , Xingjian Ma , Shirui Bian , Yichi Zhang , Xi Chen , Wei Lin

After a more than decade-long period of relatively little research activity in the area of recurrent neural networks, several new developments will be reviewed here that have allowed substantial progress both in understanding and in…

机器学习 · 计算机科学 2012-12-17 Yoshua Bengio , Nicolas Boulanger-Lewandowski , Razvan Pascanu

We study how far a diffusion process on a graph can deviate from a designed starting pattern when the pattern is generated via Laplacian regularisation. Under standard stability conditions for undirected, entrywise nonnegative graphs, we…

信号处理 · 电气工程与系统科学 2026-01-08 Ardavan Rahimian

Neural network width and depth are fundamental aspects of network topology. Universal approximation theorems provide that with increasing width or depth, there exists a neural network that approximates a function arbitrarily well. These…

机器学习 · 计算机科学 2019-10-31 Ibrohim Nosirov , Jeffrey M. Hokanson

We study the effects of mobility on two crucial characteristics in multi-scale dynamic networks: percolation and connection times. Our analysis provides insights into the question, to what extent long-time averages are well-approximated by…

概率论 · 数学 2021-03-05 Christian Hirsch , Benedikt Jahnel , Elie Cali

The NTK is a widely used tool in the theoretical analysis of deep learning, allowing us to look at supervised deep neural networks through the lenses of kernel regression. Recently, several works have investigated kernel models for…

机器学习 · 计算机科学 2025-05-06 Maximilian Fleissner , Gautham Govind Anil , Debarghya Ghoshdastidar

Recent work by Baratin et al. (2021) sheds light on an intriguing pattern that occurs during the training of deep neural networks: some layers align much more with data compared to other layers (where the alignment is defined as the…

机器学习 · 统计学 2023-04-12 Yizhang Lou , Chris Mingard , Yoonsoo Nam , Soufiane Hayou

We study free string propagation in families of plane wave geometries developing strong scale-invariant singularities in certain limits. We relate the singular limit of the evolution for all excited string modes to that of the…

高能物理 - 理论 · 物理学 2009-05-01 Ben Craps , Frederik De Roo , Oleg Evnin

Ability of deep networks to extract high level features and of recurrent networks to perform time-series inference have been studied. In view of universality of one hidden layer network at approximating functions under weak constraints, the…

神经与进化计算 · 计算机科学 2014-12-19 Sharat C. Prasad , Piyush Prasad

We consider the problem of linear fitting of noisy data in the case of broad (say $\alpha$-stable) distributions of random impacts ("noise"), which can lack even the first moment. This situation, common in statistical physics of small…

数据分析、统计与概率 · 物理学 2015-05-27 Eugene B. Postnikov , Igor M. Sokolov

A longstanding challenge for the Machine Learning community is the one of developing models that are capable of processing and learning from very long sequences of data. The outstanding results of Transformers-based networks (e.g., Large…

机器学习 · 计算机科学 2024-02-15 Matteo Tiezzi , Michele Casoni , Alessandro Betti , Tommaso Guidi , Marco Gori , Stefano Melacci

We numerically analyze the distribution of scattering resonance widths in one- and quasi-one dimensional tight binding models, in the localized regime. We detect and discuss an algebraic decay of the distribution, similar, though not…

介观与纳米尺度物理 · 物理学 2009-10-31 M. Terraneo , I. Guarneri

Finite-width fully connected neural networks with Gaussian-initialized weights deviate from their infinite-width Gaussian limit, exhibiting non-vanishing higher-order cumulants. We approximate these deviations, for a neural network…

机器学习 · 统计学 2026-05-26 Lucia Celli

Using Stein's method techniques introduced by Chatterjee (2008) and further extended by Kasprzak and Peccati (2022) and by Lachi\`eze-Rey and Peccati (2017), we derive novel quantitative bounds on the convergence in distribution of…

概率论 · 数学 2026-01-30 Lucia Celli

Common to all different kinds of recurrent neural networks (RNNs) is the intention to model relations between data points through time. When there is no immediate relationship between subsequent data points (like when the data points are…

机器学习 · 计算机科学 2022-12-22 Steffen Illium , Thore Schillman , Robert Müller , Thomas Gabor , Claudia Linnhoff-Popien

In this paper, we investigate how much of the numerical artefacts introduced by finite system size and choice of boundary conditions can be removed by finite size scaling, for strongly-correlated systems with quasi-long-range order.…

强关联电子 · 物理学 2015-05-19 Sisi Tan , Siew Ann Cheong

The evolution of a deep neural network trained by the gradient descent can be described by its neural tangent kernel (NTK) as introduced in [20], where it was proven that in the infinite width limit the NTK converges to an explicit limiting…

机器学习 · 计算机科学 2019-09-19 Jiaoyang Huang , Horng-Tzer Yau