English
Related papers

Related papers: Depth Separations in Neural Networks: What is Actu…

200 papers

In this paper, we provide a theoretical analysis of the inductive biases in convolutional neural networks (CNNs). We start by examining the universality of CNNs, i.e., the ability to approximate any continuous functions. We prove that a…

Machine Learning · Computer Science 2024-01-23 Zihao Wang , Lei Wu

We study the complexity of optimizing nonsmooth nonconvex Lipschitz functions by producing $(\delta,\epsilon)$-stationary points. Several recent works have presented randomized algorithms that produce such points using $\tilde…

Machine Learning · Computer Science 2025-05-05 Michael I. Jordan , Guy Kornowski , Tianyi Lin , Ohad Shamir , Manolis Zampetakis

It is a highly desirable property for deep networks to be robust against small input changes. One popular way to achieve this property is by designing networks with a small Lipschitz constant. In this work, we propose a new technique for…

Machine Learning · Computer Science 2023-09-04 Bernd Prach , Christoph H. Lampert

This paper demonstrates that when a shallow neural network with a Lipschitz continuous activation function is trained using either empirical or population risk to approximate a target function that is $r$ times continuously differentiable…

Machine Learning · Computer Science 2026-03-06 Sanghoon Na , Haizhao Yang

We study expressive power of shallow and deep neural networks with piece-wise linear activation functions. We establish new rigorous upper and lower bounds for the network complexity in the setting of approximations in Sobolev spaces. In…

Machine Learning · Computer Science 2017-05-02 Dmitry Yarotsky

We consider Kolmogorov widths of finite sets of functions. Any orthonormal system of $N$ functions is rigid in $L_2$, i.e. it cannot be well approximated by linear subspaces of dimension essentially smaller than $N$. This is not true for…

Functional Analysis · Mathematics 2024-01-30 Yuri Malykhin

We investigate the approximation capabilities of dense neural networks. While universal approximation theorems establish that sufficiently large architectures can approximate arbitrary continuous functions if there are no restrictions on…

Machine Learning · Computer Science 2026-05-19 Levi Rauchwerger , Stefanie Jegelka , Ron Levie

Lipschitz extensions were recently proposed as a tool for designing node differentially private algorithms. However, efficiently computable Lipschitz extensions were known only for 1-dimensional functions (that is, functions that output a…

Cryptography and Security · Computer Science 2015-04-30 Sofya Raskhodnikova , Adam Smith

A key challenge in scientific machine learning is solving partial differential equations (PDEs) on complex domains, where the curved geometry complicates the approximation of functions and their derivatives required by differential…

Numerical Analysis · Mathematics 2025-09-26 Hanfei Zhou , Lei Shi

A deep approximation is an approximating function defined by composing more than one layer of simple functions. We study deep approximations of functions of one variable using layers consisting of low-degree polynomials or simple conformal…

Numerical Analysis · Mathematics 2025-04-25 Kingsley Yeon

We demonstrate two new important properties of the 1-path-norm of shallow neural networks. First, despite its non-smoothness and non-convexity it allows a closed form proximal operator which can be efficiently computed, allowing the use of…

Machine Learning · Computer Science 2020-07-16 Fabian Latorre , Paul Rolland , Nadav Hallak , Volkan Cevher

Residual neural networks are state-of-the-art deep learning models. Their continuous-depth analog, neural ordinary differential equations (ODEs), are also widely used. Despite their success, the link between the discrete and continuous…

Machine Learning · Statistics 2024-07-08 Pierre Marion , Yu-Han Wu , Michael E. Sander , Gérard Biau

Single hidden layer feedforward neural networks can represent multivariate functions that are sums of ridge functions. These ridge functions are defined via an activation function and customizable weights. The paper deals with best…

Functional Analysis · Mathematics 2020-11-24 Steffen Goebbels

As demonstrated in many areas of real-life applications, neural networks have the capability of dealing with high dimensional data. In the fields of optimal control and dynamical systems, the same capability was studied and verified in many…

Machine Learning · Computer Science 2020-12-04 Wei Kang , Qi Gong

We show the existence of a deep neural network capable of approximating a wide class of high-dimensional approximations. The construction of the proposed neural network is based on a quasi-optimal polynomial approximation. We show that this…

Numerical Analysis · Mathematics 2019-12-09 Joseph Daws , Clayton Webster

The standard Universal Approximation Theorem for operator neural networks (NNs) holds for arbitrary width and bounded depth. Here, we prove that operator NNs of bounded width and arbitrary depth are universal approximators for continuous…

Machine Learning · Computer Science 2021-09-24 Annan Yu , Chloé Becquey , Diana Halikias , Matthew Esmaili Mallory , Alex Townsend

Lipschitz constants of neural networks allow for guarantees of robustness in image classification, safety in controller design, and generalizability beyond the training data. As calculating Lipschitz constants is NP-hard, techniques for…

Machine Learning · Computer Science 2024-01-09 Anton Xue , Lars Lindemann , Alexander Robey , Hamed Hassani , George J. Pappas , Rajeev Alur

We study the approximation of the median of $d$ inputs using ReLU neural networks. We present depth-width tradeoffs under several settings, culminating in a constant-depth, linear-width construction that achieves exponentially small…

Machine Learning · Computer Science 2026-02-10 Abhigyan Dutta , Itay Safran , Paul Valiant

Neural operator (NO) architectures learn nonlinear maps between infinite-dimensional function spaces and are widely used to accelerate simulation and enable data-driven model discovery. While universality results ensure expressivity, they…

Optimization and Control · Mathematics 2026-03-02 Takashi Furuya , Anastasis Kratsios

Error bounds and complexity bounds in numerical analysis and information-based complexity are often proved for functions that are defined on very simple domains, such as a cube, a torus, or a sphere. We study optimal error bounds for the…

Numerical Analysis · Mathematics 2020-01-15 Erich Novak