Related papers: Representing smooth functions as compositions of n…
This paper establishes the (nearly) optimal approximation error characterization of deep rectified linear unit (ReLU) networks for smooth functions in terms of both width and depth simultaneously. To that end, we first prove that…
In this paper, we analyze the mirror descent algorithm for non-smooth optimization problems in which the objective function is relatively strongly convex, without relying on the standard Lipschitz continuity assumption commonly used in the…
In this paper, we discuss the problem of minimizing the sum of two convex functions: a smooth function plus a non-smooth function. Further, the smooth part can be expressed by the average of a large number of smooth component functions, and…
Let $0 \leq \alpha<n$, $M_{\alpha}$ be the fractional maximal operator, $M^{\sharp}$ be the sharp maximal operator and $b$ be the locally integrable function. Denote by $[b, M_{\alpha}]$ and $[b, M^{\sharp}]$ be the commutators of the…
Invertible neural networks (INNs) represent an important class of deep neural network architectures that have been widely used in several applications. The universal approximation properties of INNs have also been established recently.…
The paper presents a new descent algorithm for locally Lipschitz continuous functions $f:X\to\mathbb{R}$. The selection of a descent direction at some iteration point $x$ combines an approximation of the set-valued gradient of $f$ on a…
In this work, we propose a dissipativity-based method for Lipschitz constant estimation of 1D convolutional neural networks (CNNs). In particular, we analyze the dissipativity properties of convolutional, pooling, and fully connected layers…
We prove that any continuous function f from [0,1]^n to R representable by a finite computation tree with N internal nodes and compositional sparsity s = O(1) admits a deep Kolmogorov-Arnold Network (KAN) representation. Each internal node…
In this paper, we establish a comprehensive characterization of the generalized Lipschitz classes through the study of the rate of convergence of a family of semi-discrete sampling operators, of Durrmeyer type, in $L^p$-setting. To achieve…
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer,…
Lattice Lipschitz operators define a new class of nonlinear Banach-lattice-valued maps that can be written as diagonal functions with respect to a certain basis. In the $n-$dimensional case, such a map can be represented as a vector of size…
Let (X, d) be a quasi-convex, complete and separable metric space with reference probability measure m. We prove that the set of of real valued Lipschitz function with non zero point-wise Lipschitz constant m-almost everywhere is residual,…
Given cell-average data values of a piecewise smooth bivariate function $f$ within a domain $\Omega$, we look for a piecewise adaptive approximation to $f$. We are interested in an explicit and global (smooth) approach. Bivariate…
Many problems of theoretical and practical interest involve finding an optimum over a family of convex functions. For instance, finding the projection on the convex functions in $H^k(\Omega)$, and optimizing functionals arising from some…
Preconditioning is a crucial operation in gradient-based numerical optimisation. It helps decrease the local condition number of a function by appropriately transforming its gradient. For a convex function, where the gradient can be…
We show that gradient descent can converge to any local minimum of a smooth semi-algebraic function. This holds if the step sizes are nonsummable and sufficiently small. The same results hold for the subgradient method on locally Lipschitz…
This paper presents an extension of stochastic gradient descent for the minimization of Lipschitz continuous loss functions. Our motivation is for use in non-smooth non-convex stochastic optimization problems, which are frequently…
To solve convex optimization problems with a noisy gradient input, we analyze the global behavior of subgradient-like flows under stochastic errors. The objective function is composite, being equal to the sum of two convex functions, one…
We show that finite-width deep ReLU neural networks yield rate-distortion optimal approximation (B\"olcskei et al., 2018) of polynomials, windowed sinusoidal functions, one-dimensional oscillatory textures, and the Weierstrass function, a…
We prove weak and strong maximum principles, including a Hopf lemma, for smooth subsolutions to equations defined by linear, second-order, partial differential operators whose principal symbols vanish along a portion of the domain boundary.…