English
Related papers

Related papers: Wide neural networks with general weights: converg…

200 papers

We propose probabilistic representations for inverse Stein operators (i.e. solutions to Stein equations) under general conditions; in particular we deduce new simple expressions for the Stein kernel. These representations allow to deduce…

Probability · Mathematics 2019-06-21 Marie Ernst , Gesine Reinert , Yvik Swan

We combine Stein's method with Malliavin calculus in order to obtain explicit bounds in the multidimensional normal approximation (in the Wasserstein distance) of functionals of Gaussian fields. Our results generalize and refine the main…

Probability · Mathematics 2008-11-19 Ivan Nourdin , Giovanni Peccati , Anthony Réveillac

Fully-connected deep neural networks with weights initialized from independent Gaussian distributions can be tuned to criticality, which prevents the exponential growth or decay of signals propagating through the network. However, such…

Machine Learning · Computer Science 2024-06-13 Hannah Day , Yonatan Kahn , Daniel A. Roberts

We study the discrepancy between the distribution of a vector-valued functional of i.i.d. random elements and that of a Gaussian vector. Our main contribution is an explicit bound on the convex distance between the two distributions,…

Probability · Mathematics 2022-03-25 Mikołaj J. Kasprzak , Giovanni Peccati

Convolutional neural networks (CNNs) have achieved breakthrough performances in a wide range of applications including image classification, semantic segmentation, and object detection. Previous research on characterizing the generalization…

Machine Learning · Statistics 2019-10-04 Shan Lin , Jingwei Zhang

In a recent paper, Gaunt 2020 extended Stein's method to limit distributions that can be represented as a function $g:\mathbb{R}^d\rightarrow\mathbb{R}$ of a centered multivariate normal random vector $\Sigma^{1/2}\mathbf{Z}$ with…

Probability · Mathematics 2022-09-21 Robert E. Gaunt , Heather Sutcliffe

This paper is motivated by structured sparsity for deep neural network training. We study a weighted group L0-norm constraint, and present the projection and normal cone of this set. Using randomized smoothing, we develop zeroth and…

Optimization and Control · Mathematics 2022-12-22 Michael R. Metel

Overparametrization is a key factor in the absence of convexity to explain global convergence of gradient descent (GD) for neural networks. Beside the well studied lazy regime, infinite width (mean field) analysis has been developed for…

Neural and Evolutionary Computing · Computer Science 2023-02-07 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

How much information does a learning algorithm extract from the training data and store in a neural network's weights? Too much, and the network would overfit to the training data. Too little, and the network would not fit to anything at…

Machine Learning · Computer Science 2021-03-02 Jeremy Bernstein , Yisong Yue

This work analyzes Graph Neural Networks, a generalization of Fully-Connected Deep Neural Nets on Graph structured data, when their width, that is the number of nodes in each fullyconnected layer is increasing to infinity. Infinite Width…

Machine Learning · Computer Science 2023-11-21 Yunus Cobanoglu

Algorithm-dependent generalization error bounds are central to statistical learning theory. A learning algorithm may use a large hypothesis space, but the limited number of iterations controls its model capacity and generalization error.…

Machine Learning · Computer Science 2017-07-20 Wenlong Mou , Liwei Wang , Xiyu Zhai , Kai Zheng

Covering numbers of (deep) ReLU networks have been used to characterize approximation-theoretic performance, to upper-bound prediction error in nonparametric regression, and to quantify classification capacity. These results rely on…

Machine Learning · Statistics 2026-03-04 Weigutian Ou , Helmut Bölcskei

The limit of infinite width allows for substantial simplifications in the analytical study of over-parameterised neural networks. With a suitable random initialisation, an extremely large network exhibits an approximately Gaussian…

Machine Learning · Statistics 2023-02-14 Eugenio Clerico , George Deligiannidis , Arnaud Doucet

From the classical and influential works of Neal (1996), it is known that the infinite width scaling limit of a Bayesian neural network with one hidden layer is a Gaussian process, when the network weights have bounded prior variance.…

Machine Learning · Statistics 2024-06-06 Jorge Loría , Anindya Bhadra

We introduce so-called functional input neural networks defined on a possibly infinite dimensional weighted space with values also in a possibly infinite dimensional output space. To this end, we use an additive family to map the input…

Machine Learning · Statistics 2025-12-03 Christa Cuchiero , Philipp Schmocker , Josef Teichmann

Given a reference random variable, we study the solution of its Stein equation and obtain universal bounds on its first and second derivatives. We then extend the analysis of Nourdin and Peccati by bounding the Fortet-Mourier and…

Probability · Mathematics 2017-12-13 Richard Eden , Juan Víquez

In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime'…

Machine Learning · Computer Science 2025-05-09 Qiyang Han , Masaaki Imaizumi

We analyze single-layer neural networks with the Xavier initialization in the asymptotic regime of large numbers of hidden units and large numbers of stochastic gradient descent training steps. The evolution of the neural network during…

Probability · Mathematics 2022-04-13 Justin Sirignano , Konstantinos Spiliopoulos

This work aims to provide understandings on the remarkable success of deep convolutional neural networks (CNNs) by theoretically analyzing their generalization performance and establishing optimization guarantees for gradient descent based…

Machine Learning · Computer Science 2018-05-29 Pan Zhou , Jiashi Feng

This work develops a mean-field analysis for the asymptotic behavior of deep BitNet-like architectures as smooth quantization parameters approach zero. We establish that empirical measures of latent weights converge weakly to solutions of…

Optimization and Control · Mathematics 2025-09-03 Dongwon Kim , Dongseok Lee
‹ Prev 1 4 5 6 7 8 10 Next ›