English
Related papers

Related papers: Infinite-Width Limit of a Single Attention Layer: …

200 papers

We study approximation limits of single-hidden-layer neural networks with analytic activation functions under global coefficient constraints. Under uniform $\ell^1$ bounds, or more generally sub-exponential growth of the coefficients, we…

Mathematical Finance · Quantitative Finance 2026-01-09 Jean-Gabriel Attali

Two aspects of neural networks that have been extensively studied in the recent literature are their function approximation properties and their training by gradient descent methods. The approximation problem seeks accurate approximations…

Machine Learning · Computer Science 2022-09-20 R. Gentile , G. Welper

We carry out an information-theoretical analysis of a two-layer neural network trained from input-output pairs generated by a teacher network with matching architecture, in overparametrized regimes. Our results come in the form of bounds…

Machine Learning · Computer Science 2023-07-13 Francesco Camilli , Daria Tieplova , Jean Barbier

We establish the universal approximation capability of single-layer, single-head self- and cross-attention mechanisms with minimal attached structures. Our key insight is to interpret single-head attention as an input domain-partition…

Machine Learning · Computer Science 2025-04-29 Hude Liu , Jerry Yao-Chieh Hu , Zhao Song , Han Liu

We establish central and non-central limit theorems for sequences of functionals of the Gaussian output of an infinitely-wide random neural network on the d-dimensional sphere . We show that the asymptotic behaviour of these functionals as…

Probability · Mathematics 2026-04-24 Simmaco Di Lillo , Leonardo Maini , Domenico Marinucci

This work studies approximation based on single-hidden-layer feedforward and recurrent neural networks with randomly generated internal weights. These methods, in which only the last layer of weights and a few hyperparameters are optimized,…

Probability · Mathematics 2021-02-17 Lukas Gonon , Lyudmila Grigoryeva , Juan-Pablo Ortega

Sparse neural networks promise efficiency, yet training them effectively remains a fundamental challenge. Despite advances in pruning methods that create sparse architectures, understanding why some sparse structures are better trainable…

Machine Learning · Computer Science 2025-10-21 Hoang Pham , The-Anh Ta , Tom Jacobs , Rebekka Burkholz , Long Tran-Thanh

The analytic inference, e.g. predictive distribution being in closed form, may be an appealing benefit for machine learning practitioners when they treat wide neural networks as Gaussian process in Bayesian setting. The realistic widths,…

Disordered Systems and Neural Networks · Physics 2023-08-01 Chi-Ken Lu

Overparametrization is a key factor in the absence of convexity to explain global convergence of gradient descent (GD) for neural networks. Beside the well studied lazy regime, infinite width (mean field) analysis has been developed for…

Neural and Evolutionary Computing · Computer Science 2023-02-07 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

We consider fully connected feed-forward deep neural networks (NNs) where weights and biases are independent and identically distributed as symmetric centered stable distributions. Then, we show that the infinite wide limit of the NN, under…

Machine Learning · Statistics 2020-03-03 Stefano Favaro , Sandra Fortini , Stefano Peluchetti

By classifying infinite-width neural networks and identifying the *optimal* limit, Tensor Programs IV and V demonstrated a universal way, called $\mu$P, for *widthwise hyperparameter transfer*, i.e., predicting optimal hyperparameters of…

Neural and Evolutionary Computing · Computer Science 2023-10-13 Greg Yang , Dingli Yu , Chen Zhu , Soufiane Hayou

Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript,…

Machine Learning · Statistics 2025-06-26 Federico Bassetti , Marco Gherardi , Alessandro Ingrosso , Mauro Pastore , Pietro Rotondo

We study the expressivity of sparse maxout networks, where each neuron takes a fixed number of inputs from the previous layer and employs a, possibly multi-argument, maxout activation. This setting captures key characteristics of…

Machine Learning · Computer Science 2025-10-17 Moritz Grillo , Tobias Hofmann

The quadratic cost of softmax attention limits Transformer scalability in high-resolution vision. We introduce Infinite Self-Attention (InfSA), a spectral reformulation that treats each attention layer as a diffusion step on a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Giorgio Roffo , Hazem Abdelkawy , Nilli Lavie , Luke Palmer

For almost 70 years, researchers have typically selected the width of neural networks' layers either manually or through automated hyperparameter tuning methods such as grid search and, more recently, neural architecture search. This paper…

Machine Learning · Computer Science 2026-02-17 Federico Errica , Henrik Christiansen , Viktor Zaverkin , Mathias Niepert , Francesco Alesiani

Deep neural networks (DNNs) in the infinite width/channel limit have received much attention recently, as they provide a clear analytical window to deep learning via mappings to Gaussian Processes (GPs). Despite its theoretical appeal, this…

Machine Learning · Computer Science 2021-06-09 Gadi Naveh , Zohar Ringel

In this work, we analyze various scaling limits of the training dynamics of transformer models in the feature learning regime. We identify the set of parameterizations that admit well-defined infinite width and depth limits, allowing the…

Machine Learning · Statistics 2024-10-07 Blake Bordelon , Hamza Tahir Chaudhry , Cengiz Pehlevan

Using Stein's method techniques introduced by Chatterjee (2008) and further extended by Kasprzak and Peccati (2022) and by Lachi\`eze-Rey and Peccati (2017), we derive novel quantitative bounds on the convergence in distribution of…

Probability · Mathematics 2026-01-30 Lucia Celli

It has long been known that a single-layer fully-connected neural network with an i.i.d. prior over its parameters is equivalent to a Gaussian process (GP), in the limit of infinite network width. This correspondence enables exact Bayesian…

Wide neural networks have proven to be a rich class of architectures for both theory and practice. Motivated by the observation that finite width convolutional networks appear to outperform infinite width networks, we study scaling laws for…

Machine Learning · Computer Science 2020-08-21 Anders Andreassen , Ethan Dyer