中文
相关论文

相关论文: Neural networks: deep, shallow, or in between?

200 篇论文

Using Stein's method techniques introduced by Chatterjee (2008) and further extended by Kasprzak and Peccati (2022) and by Lachi\`eze-Rey and Peccati (2017), we derive novel quantitative bounds on the convergence in distribution of…

概率论 · 数学 2026-01-30 Lucia Celli

We obtain Lebesgue-type inequalities for the greedy algorithm for arbitrary complete seminormalized biorthogonal systems in Banach spaces. The bounds are given only in terms of the upper democracy functions of the basis and its dual. We…

泛函分析 · 数学 2017-09-18 P. M. Berná , O. Blasco , G. Garrigós , E. Hernández , T. Oikhberg

In this paper, we study approximation properties of single hidden layer neural networks with weights varying on finitely many directions and thresholds from an open interval. We obtain a necessary and at the same time sufficient measure…

机器学习 · 计算机科学 2023-04-05 Vugar Ismailov , Ekrem Savas

We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…

机器学习 · 计算机科学 2024-10-11 Semih Cayci , Atilla Eryilmaz

We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer,…

机器学习 · 计算机科学 2018-07-25 Thomas Laurent , James von Brecht

Conventional wisdom in deep learning states that increasing depth improves expressiveness but complicates optimization. This paper suggests that, sometimes, increasing depth can speed up optimization. The effect of depth on optimization is…

机器学习 · 计算机科学 2018-06-12 Sanjeev Arora , Nadav Cohen , Elad Hazan

A key attribute that drives the unprecedented success of modern Recurrent Neural Networks (RNNs) on learning tasks which involve sequential data, is their ability to model intricate long-term temporal dependencies. However, a well…

机器学习 · 计算机科学 2020-03-24 Alon Ziv

We initiate the study of the inherent tradeoffs between the size of a neural network and its robustness, as measured by its Lipschitz constant. We make a precise conjecture that, for any Lipschitz activation function and for most datasets,…

机器学习 · 计算机科学 2020-11-26 Sébastien Bubeck , Yuanzhi Li , Dheeraj Nagaraj

We develop Banach spaces for ReLU neural networks of finite depth $L$ and infinite width. The spaces contain all finite fully connected $L$-layer networks and their $L^2$-limiting objects under bounds on the natural path-norm. Under this…

机器学习 · 统计学 2020-07-31 Weinan E , Stephan Wojtowytsch

This paper investigates achievable information rates and error exponents of mismatched decoding when the channel belongs to the class of channels that are close to the decoding metric in terms of relative entropy. For both discrete- and…

信息论 · 计算机科学 2025-05-28 Priyanka Patel , Francesc Molina , Albert Guillén i Fàbregas

We define "decision swap regret" which generalizes both prediction for downstream swap regret and omniprediction, and give algorithms for obtaining it for arbitrary multi-dimensional Lipschitz loss functions in online adversarial settings.…

机器学习 · 计算机科学 2025-02-19 Jiuyao Lu , Aaron Roth , Mirah Shi

We introduce the notions of almost Lipschitz embeddability and nearly isometric embeddability. We prove that for $p\in [1,\infty]$, every proper subset of $L_p$ is almost Lipschitzly embeddable into a Banach space $X$ if and only if $X$…

度量几何 · 数学 2017-09-27 Florent Baudier , Gilles Lancien

It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of…

机器学习 · 计算机科学 2025-11-26 Antoine Ledent , Rodrigo Alves , Yunwen Lei

This paper investigates the approximation properties of shallow neural networks with activation functions that are powers of exponential functions. It focuses on the dependence of the approximation rate on the dimension and the smoothness…

机器学习 · 计算机科学 2025-10-22 Jian Lu , Xiaohuang Huang

The classical Universal Approximation Theorem holds for neural networks of arbitrary width and bounded depth. Here we consider the natural `dual' scenario for networks of bounded width and arbitrary depth. Precisely, let $n$ be the number…

机器学习 · 计算机科学 2020-06-09 Patrick Kidger , Terry Lyons

Deep neural networks have shown incredible performance for inference tasks in a variety of domains. Unfortunately, most current deep networks are enormous cloud-based structures that require significant storage space, which limits scaling…

信息论 · 计算机科学 2020-03-10 Sourya Basu , Lav R. Varshney

This paper studies the problem of how efficiently functions in the Sobolev spaces $\mathcal{W}^{s,q}([0,1]^d)$ and Besov spaces $\mathcal{B}^s_{q,r}([0,1]^d)$ can be approximated by deep ReLU neural networks with width $W$ and depth $L$,…

机器学习 · 统计学 2025-07-21 Yunfei Yang

It has been experimentally observed in recent years that multi-layer artificial neural networks have a surprising ability to generalize, even when trained with far more parameters than observations. Is there a theoretical basis for this?…

机器学习 · 统计学 2018-09-19 Andrew R. Barron , Jason M. Klusowski

Neural scaling laws relate loss to model size in large language models (LLMs), yet depth and width may contribute to performance differently, requiring more detailed studies. Here, we quantify how depth affects loss via analysis of LLMs and…

机器学习 · 计算机科学 2026-02-06 Yizhou Liu , Sara Kangaslahti , Ziming Liu , Jeff Gore

The purpose of this article is to develop a technique to estimate certain bounds for entropy numbers of diagonal operator on spaces of p-summable sequences for finite p greater than 1. The approximation method we develop in this direction…

泛函分析 · 数学 2022-07-08 K. P. Deepesh , V. B. Kiran Kumar