Related papers: Polynomial, trigonometric, and tropical activation…
This paper presents a new mathematical framework to analyze the loss functions of deep neural networks with ReLU functions. Furthermore, as as application of this theory, we prove that the loss functions can reconstruct the inputs of the…
The fundamental learning theory behind neural networks remains largely open. What classes of functions can neural networks actually learn? Why doesn't the trained network overfit when it is overparameterized? In this work, we prove that…
In the last decade, an active area of research has been devoted to design novel activation functions that are able to help deep neural networks to converge, obtaining better performance. The training procedure of these architectures usually…
This paper presents a novel method for polynomial approximation (Hermite approximation) using the fusion of value and derivative information. Therefore, the least-squares error in both domains is simultaneously minimized. A covariance…
The effective formulas reducing the two-dimensional Hermite polynomials to the classical (one-dimensional) orthogonal polynomials are given. New one-parameter generating functions for these polynomials are derived. Asymptotical formulas for…
The active bijection for oriented matroids (and real hyperplane arrangements, and graphs, as particular cases) is introduced and investigated by the authors in a series of papers. Given any oriented matroid defined on a linearly ordered…
The training of two-layer neural networks with nonlinear activation functions is an important non-convex optimization problem with numerous applications and promising performance in layerwise deep learning. In this paper, we develop exact…
We explore the phase diagram of approximation rates for deep neural networks and prove several new theoretical results. In particular, we generalize the existing result on the existence of deep discontinuous phase in ReLU networks to…
Low bit-width weights and activations are an effective way of combating the increasing need for both memory and compute power of Deep Neural Networks. In this work, we present a probabilistic training method for Neural Network with both…
We introduce a new class of non-linear models for functional data based on neural networks. Deep learning has been very successful in non-linear modeling, but there has been little work done in the functional data setting. We propose two…
Tropical refined invariants of toric surfaces constitute a fascinating interpolation between real and complex enumerative geometries via tropical geometry. They were originally introduced by Block and G\"ottsche, and further extended by…
We consider a quaternionic analogue of the univariate complex Hermite polynomials and study some of their analytic properties in some detail. We obtain their integral representation as well as the operational formulas of exponential and…
We present the soft exponential activation function for artificial neural networks that continuously interpolates between logarithmic, linear, and exponential functions. This activation function is simple, differentiable, and parameterized…
Activation functions play a decisive role in determining the capacity of Deep Neural Networks as they enable neural networks to capture inherent nonlinearities present in data fed to them. The prior research on activation functions…
In this work we introduce a new algebra of tempered generalized functions. The tempered distributions are embedded in this algebra via their Hermite expansions. The Fourier transform is naturally extended to this algebra in such a way that…
The mathematical complexity and high dimensionality of neural networks slow both training and deployment, demanding heavy computational resources. This has driven the search for alternative architectures built from novel components,…
Artificial neural networks (ANN), typically referred to as neural networks, are a class of Machine Learning algorithms and have achieved widespread success, having been inspired by the biological structure of the human brain. Neural…
Asymptotic properties of matrices are, in general, difficult to analyze with classical mathematical techniques. In very specific cases, there is a well-known connection between the asymptotic behavior of a matrix's leading eigenvector and…
The article is devoted to a new proof of the expansion for iterated Ito stochastic integrals with respect to the components of a multidimensional Wiener process. The above expansion is based on Hermite polynomials and generalized multiple…
Gaussian radial basis functions can be an accurate basis for multivariate interpolation. In practise, high accuracies are often achieved in the flat limit where the interpolation matrix becomes increasingly ill-conditioned. Stable…