Related papers: Large deviation principles for convolutional Bayes…
Recent works show an intriguing phenomenon of Frequency Principle (F-Principle) that deep neural networks (DNNs) fit the target function from low to high frequency during the training, which provides insight into the training and…
We study an inhomogeneous sparse random graph on [N] = {1, . . . , N } as introduced in a seminal paper by Bollobas, Janson and Riordan (2007): vertices have a type (here in a compact metric space S), and edges between different vertices…
Despite exceptional predictive performance of Deep sequence models (DSMs), the main concern of their deployment centers around the lack of uncertainty awareness. In contrast, probabilistic models quantify the uncertainty associated with…
Generalized Large deviation principles was developed for Colombeau-Ito SDE with a random coefficients. We is significantly expand the classical theory of large deviations for randomly perturbed dynamical systems developed by Freidlin and…
For an arbitrary negative Schwarzian unimodal map with non-flat critical point, we establish the level-2 Large Deviation Principle (LDP) for empirical distributions. We also give an example of a multimodal map for which the level-2 LDP does…
We introduce a principled approach for unsupervised structure learning of deep neural networks. We propose a new interpretation for depth and inter-layer connectivity where conditional independencies in the input distribution are encoded…
This article studies the infinite-width limit of deep feedforward neural networks whose weights are dependent, and modelled via a mixture of Gaussian distributions. Each hidden node of the network is assigned a nonnegative random variable…
In this paper, we rigorously derive Central Limit Theorems (CLT) for Bayesian two-layerneural networks in the infinite-width limit and trained by variational inference on a regression task. The different networks are trained via different…
Bayesian neural networks (BNNs) have been long considered an ideal, yet unscalable solution for improving the robustness and the predictive uncertainty of deep neural networks. While they could capture more accurately the posterior…
Neural Processes (NPs) are meta-learning models that learn to map sets of observations to approximations of the corresponding posterior predictive distributions. By accommodating variable-sized, unstructured collections of observations and…
Following the traditional paradigm of convolutional neural networks (CNNs), modern CNNs manage to keep pace with more recent, for example transformer-based, models by not only increasing model depth and width but also the kernel size. This…
We show generalisation error bounds for deep learning with two main improvements over the state of the art. (1) Our bounds have no explicit dependence on the number of classes except for logarithmic factors. This holds even when formulating…
This work concerns about multiscale multivalued McKean-Vlasov stochastic systems. First of all, we use a contractive mapping principle to establish the well-posedness for fully coupled multivalued McKean-Vlasov stochastic systems under…
We establish a functional large deviation principle for fully connected multi-layer perceptrons with i.i.d. Gaussian weights (LeCun initialization) and general Lipschitz activation functions, including therefore the popular case of ReLU.…
A key property of neural networks driving their success is their ability to learn features from data. Understanding feature learning from a theoretical viewpoint is an emerging field with many open questions. In this work we capture…
We consider a family of positive operator valued measures associated with representations of compact connected Lie groups. For many independent copies of a single state and a tensor power representation we show that the observed probability…
In this article we obtain large deviation asymptotics for supercritical communication networks modelled as signal-interference-noise ratio networks. To do this, we define the empirical power measure and the empirical connectivity measure,…
In decision-making systems, it is important to have classifiers that have calibrated uncertainties, with an optimisation objective that can be used for automated model selection and training. Gaussian processes (GPs) provide uncertainty…
The main results in this paper concern large deviations for families of non-Gaussian processes obtained as suitable perturbations of continuous centered multivariate Gaussian processes which satisfy a large deviation principle. We present…
Neural network approaches for meta-learning distributions over functions have desirable properties such as increased flexibility and a reduced complexity of inference. Building on the successes of denoising diffusion models for generative…