中文
相关论文

相关论文: Wide neural networks with general weights: converg…

200 篇论文

The analytic inference, e.g. predictive distribution being in closed form, may be an appealing benefit for machine learning practitioners when they treat wide neural networks as Gaussian process in Bayesian setting. The realistic widths,…

无序系统与神经网络 · 物理学 2023-08-01 Chi-Ken Lu

Bayesian neural networks (BNNs) combine the expressive power of deep learning with the advantages of Bayesian formalism. In recent years, the analysis of wide, deep BNNs has provided theoretical insight into their priors and posteriors.…

机器学习 · 计算机科学 2022-02-24 Beau Coker , Wessel P. Bruinsma , David R. Burt , Weiwei Pan , Finale Doshi-Velez

This paper is devoted to the estimation of the Lipschitz constant of general neural network architectures using semidefinite programming. For this purpose, we interpret neural networks as time-varying dynamical systems, where the $k$-th…

机器学习 · 计算机科学 2024-11-26 Patricia Pauli , Dennis Gramlich , Frank Allgöwer

Recent advances in adversarial attacks and Wasserstein GANs have advocated for use of neural networks with restricted Lipschitz constants. Motivated by these observations, we study the recently introduced GroupSort neural networks, with…

机器学习 · 统计学 2021-02-09 Ugo Tanielian , Maxime Sangnier , Gerard Biau

Deep neural networks often generalize well despite heavy over-parameterization, challenging classical parameter-based analyses. We study generalization from a representation-centric perspective and analyze how the geometry of learned…

机器学习 · 计算机科学 2026-02-02 Junjie Yu , Zhuoli Ouyang , Haotian Deng , Chen Wei , Wenxiao Ma , Jianyu Zhang , Zihan Deng , Quanying Liu

We study the problem of approximately recovering a probability distribution given noisy measurements of its Chebyshev polynomial moments. This problem arises broadly across algorithms, statistics, and machine learning. By leveraging a…

数据结构与算法 · 计算机科学 2026-05-20 Cameron Musco , Christopher Musco , Lucas Rosenblatt , Apoorv Vikram Singh

We combine Malliavin calculus with Stein's method, in order to derive explicit bounds in the Gaussian and Gamma approximations of random variables in a fixed Wiener chaos of a general Gaussian process. We also prove results concerning…

概率论 · 数学 2008-05-10 Ivan Nourdin , Giovanni Peccati

Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in…

机器学习 · 统计学 2024-11-01 Cheng Gao , Yuan Cao , Zihao Li , Yihan He , Mengdi Wang , Han Liu , Jason Matthew Klusowski , Jianqing Fan

This paper advances the state of the art in girth approximation within the CONGEST model. Manoharan and Ramachandran [PODC '24] provided the first significant improvement in girth approximation in over a decade. We build on this momentum…

数据结构与算法 · 计算机科学 2026-03-31 Shiri Chechik , Gur Lifshitz , Doron Mukhtar

Obtaining sharp Lipschitz constants for feed-forward neural networks is essential to assess their robustness in the face of perturbations of their inputs. We derive such constants in the context of a general layered network model involving…

最优化与控制 · 数学 2020-06-23 Patrick L. Combettes , Jean-Christophe Pesquet

We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training $L$-hidden-layer linear residual networks (ResNets). We prove that for training deep residual networks with certain linear transformations…

机器学习 · 计算机科学 2020-03-03 Difan Zou , Philip M. Long , Quanquan Gu

In this paper, we present BERN-NN as an efficient tool to perform bound propagation of Neural Networks (NNs). Bound propagation is a critical step in wide range of NN model checkers and reachability analysis tools. Given a bounded input…

机器学习 · 计算机科学 2022-11-29 Wael Fatnassi , Haitham Khedr , Valen Yamamoto , Yasser Shoukry

We present an explicit deep neural network construction that transforms uniformly distributed one-dimensional noise into an arbitrarily close approximation of any two-dimensional Lipschitz-continuous target distribution. The key ingredient…

机器学习 · 计算机科学 2021-06-08 Dmytro Perekrestenko , Stephan Müller , Helmut Bölcskei

This paper derives non-asymptotic error bounds for nonlinear stochastic approximation algorithms in the Wasserstein-$p$ distance. To obtain explicit finite-sample guarantees for the last iterate, we develop a coupling argument that compares…

机器学习 · 计算机科学 2026-02-03 Seo Taek Kong , R. Srikant

Weight sharing, equivariance, and local filters, as in convolutional neural networks, are believed to contribute to the sample efficiency of neural networks. However, it is not clear how each one of these design choices contributes to the…

机器学习 · 计算机科学 2025-01-27 Arash Behboodi , Gabriele Cesa

We analyze initialization dynamics for LDLT-based $\mathcal{L}$-Lipschitz layers by deriving the exact marginal output variance when the underlying parameter matrix $W_0\in \mathbb{R}^{m\times n}$ is initialized with IID Gaussian entries…

机器学习 · 计算机科学 2026-01-14 Marius F. R. Juston , Ramavarapu S. Sreenivas , Dustin Nottage , Ahmet Soylemezoglu

Understanding capabilities and limitations of different network architectures is of fundamental importance to machine learning. Bayesian inference on Gaussian processes has proven to be a viable approach for studying recurrent and deep…

无序系统与神经网络 · 物理学 2022-10-17 Kai Segadlo , Bastian Epping , Alexander van Meegen , David Dahmen , Michael Krämer , Moritz Helias

Recent advances in large-margin classification of data residing in general metric spaces (rather than Hilbert spaces) enable classification under various natural metrics, such as string edit and earthmover distance. A general framework…

机器学习 · 计算机科学 2014-07-14 Lee-Ad Gottlieb , Aryeh Kontorovich , Robert Krauthgamer

Estimating the Lipschitz constant of deep neural networks is of growing interest as it is useful for informing on generalisability and adversarial robustness. Convolutional neural networks (CNNs) in particular, underpin much of the recent…

机器学习 · 计算机科学 2024-08-08 Yusuf Sulehman , Tingting Mu

In this paper we study the generalization capabilities of fully-connected neural networks trained in the context of time series forecasting. Time series do not satisfy the typical assumption in statistical learning theory of the data being…

机器学习 · 统计学 2019-07-30 Anastasia Borovykh , Cornelis W. Oosterlee , Sander M. Bohte