中文
相关论文

相关论文: Wide neural networks with general weights: converg…

200 篇论文

We propose probabilistic representations for inverse Stein operators (i.e. solutions to Stein equations) under general conditions; in particular we deduce new simple expressions for the Stein kernel. These representations allow to deduce…

概率论 · 数学 2019-06-21 Marie Ernst , Gesine Reinert , Yvik Swan

We combine Stein's method with Malliavin calculus in order to obtain explicit bounds in the multidimensional normal approximation (in the Wasserstein distance) of functionals of Gaussian fields. Our results generalize and refine the main…

概率论 · 数学 2008-11-19 Ivan Nourdin , Giovanni Peccati , Anthony Réveillac

Fully-connected deep neural networks with weights initialized from independent Gaussian distributions can be tuned to criticality, which prevents the exponential growth or decay of signals propagating through the network. However, such…

机器学习 · 计算机科学 2024-06-13 Hannah Day , Yonatan Kahn , Daniel A. Roberts

We study the discrepancy between the distribution of a vector-valued functional of i.i.d. random elements and that of a Gaussian vector. Our main contribution is an explicit bound on the convex distance between the two distributions,…

概率论 · 数学 2022-03-25 Mikołaj J. Kasprzak , Giovanni Peccati

Convolutional neural networks (CNNs) have achieved breakthrough performances in a wide range of applications including image classification, semantic segmentation, and object detection. Previous research on characterizing the generalization…

机器学习 · 统计学 2019-10-04 Shan Lin , Jingwei Zhang

In a recent paper, Gaunt 2020 extended Stein's method to limit distributions that can be represented as a function $g:\mathbb{R}^d\rightarrow\mathbb{R}$ of a centered multivariate normal random vector $\Sigma^{1/2}\mathbf{Z}$ with…

概率论 · 数学 2022-09-21 Robert E. Gaunt , Heather Sutcliffe

This paper is motivated by structured sparsity for deep neural network training. We study a weighted group L0-norm constraint, and present the projection and normal cone of this set. Using randomized smoothing, we develop zeroth and…

最优化与控制 · 数学 2022-12-22 Michael R. Metel

Overparametrization is a key factor in the absence of convexity to explain global convergence of gradient descent (GD) for neural networks. Beside the well studied lazy regime, infinite width (mean field) analysis has been developed for…

神经与进化计算 · 计算机科学 2023-02-07 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

How much information does a learning algorithm extract from the training data and store in a neural network's weights? Too much, and the network would overfit to the training data. Too little, and the network would not fit to anything at…

机器学习 · 计算机科学 2021-03-02 Jeremy Bernstein , Yisong Yue

This work analyzes Graph Neural Networks, a generalization of Fully-Connected Deep Neural Nets on Graph structured data, when their width, that is the number of nodes in each fullyconnected layer is increasing to infinity. Infinite Width…

机器学习 · 计算机科学 2023-11-21 Yunus Cobanoglu

Algorithm-dependent generalization error bounds are central to statistical learning theory. A learning algorithm may use a large hypothesis space, but the limited number of iterations controls its model capacity and generalization error.…

机器学习 · 计算机科学 2017-07-20 Wenlong Mou , Liwei Wang , Xiyu Zhai , Kai Zheng

Covering numbers of (deep) ReLU networks have been used to characterize approximation-theoretic performance, to upper-bound prediction error in nonparametric regression, and to quantify classification capacity. These results rely on…

机器学习 · 统计学 2026-03-04 Weigutian Ou , Helmut Bölcskei

The limit of infinite width allows for substantial simplifications in the analytical study of over-parameterised neural networks. With a suitable random initialisation, an extremely large network exhibits an approximately Gaussian…

机器学习 · 统计学 2023-02-14 Eugenio Clerico , George Deligiannidis , Arnaud Doucet

From the classical and influential works of Neal (1996), it is known that the infinite width scaling limit of a Bayesian neural network with one hidden layer is a Gaussian process, when the network weights have bounded prior variance.…

机器学习 · 统计学 2024-06-06 Jorge Loría , Anindya Bhadra

We introduce so-called functional input neural networks defined on a possibly infinite dimensional weighted space with values also in a possibly infinite dimensional output space. To this end, we use an additive family to map the input…

机器学习 · 统计学 2025-12-03 Christa Cuchiero , Philipp Schmocker , Josef Teichmann

Given a reference random variable, we study the solution of its Stein equation and obtain universal bounds on its first and second derivatives. We then extend the analysis of Nourdin and Peccati by bounding the Fortet-Mourier and…

概率论 · 数学 2017-12-13 Richard Eden , Juan Víquez

In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime'…

机器学习 · 计算机科学 2025-05-09 Qiyang Han , Masaaki Imaizumi

We analyze single-layer neural networks with the Xavier initialization in the asymptotic regime of large numbers of hidden units and large numbers of stochastic gradient descent training steps. The evolution of the neural network during…

概率论 · 数学 2022-04-13 Justin Sirignano , Konstantinos Spiliopoulos

This work aims to provide understandings on the remarkable success of deep convolutional neural networks (CNNs) by theoretically analyzing their generalization performance and establishing optimization guarantees for gradient descent based…

机器学习 · 计算机科学 2018-05-29 Pan Zhou , Jiashi Feng

This work develops a mean-field analysis for the asymptotic behavior of deep BitNet-like architectures as smooth quantization parameters approach zero. We establish that empirical measures of latent weights converge weakly to solutions of…

最优化与控制 · 数学 2025-09-03 Dongwon Kim , Dongseok Lee