English
Related papers

Related papers: Random Neural Networks in the Infinite Width Limit…

200 papers

We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…

Machine Learning · Computer Science 2024-10-11 Semih Cayci , Atilla Eryilmaz

Deep learning researchers commonly suggest that converged models are stuck in local minima. More recently, some researchers observed that under reasonable assumptions, the vast majority of critical points are saddle points, not true minima.…

Machine Learning · Computer Science 2016-02-25 Zachary C. Lipton

We analyze the threshold network model in which a pair of vertices with random weights are connected by an edge when the summation of the weights exceeds a threshold. We prove some convergence theorems and central limit theorems on the…

Probability · Mathematics 2007-05-23 Norio Konno , Naoki Masuda , Rahul Roy , Anish Sarkar

Multi-layer feedforward networks have been used to approximate a wide range of nonlinear functions. An important and fundamental problem is to understand the learnability of a network model through its statistical risk, or the expected…

Machine Learning · Computer Science 2022-06-28 Gen Li , Jie Ding

We consider nonlinear networks as perturbations of linear ones. Based on this approach, we present novel generalization bounds that become non-vacuous for networks that are close to being linear. The main advantage over the previous works…

Machine Learning · Computer Science 2024-07-10 Eugene Golikov

The focus of this work is the convergence of non-stationary and deep Gaussian process regression. More precisely, we follow a Bayesian approach to regression or interpolation, where the prior placed on the unknown function $f$ is a…

Statistics Theory · Mathematics 2025-03-19 Conor Osborne , Aretha L. Teckentrup

An interesting approach to analyzing neural networks that has received renewed attention is to examine the equivalent kernel of the neural network. This is based on the fact that a fully connected feedforward network with one hidden layer,…

Machine Learning · Computer Science 2018-06-04 Russell Tsuchida , Farbod Roosta-Khorasani , Marcus Gallagher

The successes of modern deep machine learning methods are founded on their ability to transform inputs across multiple layers to build good high-level representations. It is therefore critical to understand this process of representation…

Machine Learning · Statistics 2023-05-26 Adam X. Yang , Maxime Robeyns , Edward Milsom , Ben Anson , Nandi Schoots , Laurence Aitchison

interpretable, and well understood models that are routinely employed even though, as is revealed through prior and posterior predictive checks, these can poorly characterise the spatial heterogeneity in the underlying process of interest.…

Machine Learning · Statistics 2024-04-08 Andrew Zammit-Mangion , Michael D. Kaminski , Ba-Hien Tran , Maurizio Filippone , Noel Cressie

We consider the infinite-width limit of a fully connected deep neural network with general weights, and we prove quantitative general bounds on the $2$-Wasserstein distance between the network and its infinite-width Gaussian limit, under…

Probability · Mathematics 2026-05-05 Filippo Giovagnini , Sotirios Kotitsas , Marco Romito

Neural networks trained to minimize the logistic (a.k.a. cross-entropy) loss with gradient-based methods are observed to perform well in many supervised classification tasks. Towards understanding this phenomenon, we analyze the training…

Optimization and Control · Mathematics 2020-06-23 Lenaic Chizat , Francis Bach

This paper studies the approximation capacity of ReLU neural networks with norm constraint on the weights. We prove upper and lower bounds on the approximation error of these networks for smooth function classes. The lower bound is derived…

Machine Learning · Computer Science 2023-03-31 Yuling Jiao , Yang Wang , Yunfei Yang

We investigate the realizations of a random Gaussian field on a finite domain of ${\mathbb R}^d$ in the limit where a given linear functional of the field is large. We prove that if its variance is bounded, the field converges uniformly and…

Probability · Mathematics 2019-02-07 Philippe Mounaix

Bayesian neural networks attempt to combine the strong predictive performance of neural networks with formal quantification of uncertainty associated with the predictive output in the Bayesian framework. However, it remains unclear how to…

Machine Learning · Statistics 2022-01-12 Takuo Matsubara , Chris J. Oates , François-Xavier Briol

In this paper, we show that although the minimizers of cross-entropy and related classification losses are off at infinity, network weights learned by gradient flow converge in direction, with an immediate corollary that network…

Machine Learning · Computer Science 2020-10-27 Ziwei Ji , Matus Telgarsky

The threshold network model is a type of finite random graphs. In this paper, we introduce a generalized threshold network model. A pair of vertices with random weights is connected by an edge when real-valued functions of the pair of…

Probability · Mathematics 2010-10-12 Yusuke Ide , Norio Konno , Naoki Masuda

We propose a theoretical understanding of neural networks in terms of Wilsonian effective field theory. The correspondence relies on the fact that many asymptotic neural networks are drawn from Gaussian processes, the analog of…

Machine Learning · Computer Science 2021-03-16 James Halverson , Anindita Maiti , Keegan Stoner

Neural networks are nowadays highly successful despite strong hardness results. The existing hardness results focus on the network architecture, and assume that the network's weights are arbitrary. A natural approach to settle the…

Machine Learning · Computer Science 2020-10-15 Amit Daniely , Gal Vardi

Bayesian networks provide a method of representing conditional independence between random variables and computing the probability distributions associated with these random variables. In this paper, we extend Bayesian network structures to…

Artificial Intelligence · Computer Science 2013-02-21 Eric Driver , Darryl Morrell

We construct flexible likelihoods for multi-output Gaussian process models that leverage neural networks as components. We make use of sparse variational inference methods to enable scalable approximate inference for the resulting class of…

Machine Learning · Statistics 2019-06-03 Martin Jankowiak , Jacob Gardner