English
Related papers

Related papers: Generalized Euler Logarithm and its Applications i…

200 papers

We present the first method to directly use a learned continuous Lagrangian to forecast the dynamics of systems governed by partial differential equations, exploiting the inherent conservative structure to achieve stable long-range…

Machine Learning · Computer Science 2026-05-11 Lyra Zhornyak , Eric Forgoston , M. Ani Hsieh

Most modern learning problems are highly overparameterized, meaning that there are many more parameters than the number of training data points, and as a result, the training loss may have infinitely many global minima (parameter vectors…

Machine Learning · Computer Science 2019-06-11 Navid Azizan , Sahin Lale , Babak Hassibi

We introduce two-scale loss functions for use in various gradient descent algorithms applied to classification problems via deep neural networks. This new method is generic in the sense that it can be applied to a wide range of machine…

Numerical Analysis · Mathematics 2021-09-03 Leonid Berlyand , Robert Creese , Pierre-Emmanuel Jabin

We propose an approach to construction of robust non-Euclidean iterative algorithms for convex composite stochastic optimization based on truncation of stochastic gradients. For such algorithms, we establish sub-Gaussian confidence bounds…

Statistics Theory · Mathematics 2019-07-08 Anatoli Juditsky , Alexander Nazin , Arkadi Nemirovsky , Alexandre Tsybakov

We derive a differential equation that governs the evolution of the generalization gap when a deep network is trained by gradient descent. This differential equation is controlled by two quantities, a contraction factor that brings together…

Machine Learning · Computer Science 2025-10-14 Rubing Yang , Pratik Chaudhari

Bayesian inference and kernel methods are well established in machine learning. The neural network Gaussian process in particular provides a concept to investigate neural networks in the limit of infinitely wide hidden layers by using…

Disordered Systems and Neural Networks · Physics 2023-11-10 Javed Lindner , David Dahmen , Michael Krämer , Moritz Helias

The recent generalizations of Boltzmann-Gibbs statistics mathematically relies on the deformed logarithmic and exponential functions defined through some deformation parameters. In the present work, we investigate whether a deformed…

Statistical Mechanics · Physics 2009-07-24 Thomas Oikonomou , G. Baris Bagci

Over the past years, there has been significant interest in understanding the implicit bias of gradient descent optimization and its connection to the generalization properties of overparametrized neural networks. Several works observed…

Optimization and Control · Mathematics 2025-03-11 Hung-Hsu Chou , Johannes Maly , Claudio Mayrink Verdun , Bernardo Freitas Paulo da Costa , Heudson Mirandola

Understanding the inner workings of neural networks, including transformers, remains one of the most challenging puzzles in machine learning. This study introduces a novel approach by applying the principles of gauge symmetries, a key…

Machine Learning · Computer Science 2024-02-06 Koji Hashimoto , Yuji Hirono , Akiyoshi Sannai

Many scientific and geometric problems exhibit general linear symmetries, yet most equivariant neural networks are built for compact groups or simple vector features, limiting their reuse on matrix-valued data such as covariances, inertias,…

Machine Learning · Computer Science 2026-02-02 Chankyo Kim , Sicheng Zhao , Minghan Zhu , Tzu-Yuan Lin , Maani Ghaffari

We investigate the convergence properties of the EM algorithm when applied to overspecified Gaussian mixture models -- that is, when the number of components in the fitted model exceeds that of the true underlying distribution. Focusing on…

Machine Learning · Statistics 2025-06-16 Zhenisbek Assylbekov , Alan Legg , Artur Pak

The Euler numbers have been widely studied. A signed version of the Euler numbers of even subscript are given by the coefficients of the exponential generating function 1/(1+x^2/2!+x^4/4!+...). Leeming and MacLeod introduced a…

Number Theory · Mathematics 2025-01-15 Bruce E. Sagan

Gating is a key feature in modern neural networks including LSTMs, GRUs and sparsely-gated deep neural networks. The backbone of such gated networks is a mixture-of-experts layer, where several experts make regression decisions and gating…

Machine Learning · Computer Science 2020-06-19 Ashok Vardhan Makkuva , Sewoong Oh , Sreeram Kannan , Pramod Viswanath

We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets. We look at two basic questions: (a) what are the…

We present a novel class of Physics-Informed Neural Networks that is formulated based on the principles of Evidential Deep Learning, where the model incorporates uncertainty quantification by learning parameters of a higher-order…

Machine Learning · Computer Science 2025-01-28 Hai Siong Tan , Kuancheng Wang , Rafe McBeth

We revisit logistic regression and its nonlinear extensions, including multilayer feedforward neural networks, by showing that these classifiers can be viewed as converting input or higher-level features into Dempster-Shafer mass functions…

Machine Learning · Computer Science 2019-12-13 Thierry Denoeux

The data consistency for the physical forward model is crucial in inverse problems, especially in MR imaging reconstruction. The standard way is to unroll an iterative algorithm into a neural network with a forward model embedded. The…

Image and Video Processing · Electrical Eng. & Systems 2023-06-28 Guanxiong Luo , Mengmeng Kuang , Peng Cao

We define the generalized-Euler-constant function $\gamma(z)=\sum_{n=1}^{\infty} z^{n-1} (\frac{1}{n}-\log \frac{n+1}{n})$ when $|z|\leq 1$. Its values include both Euler's constant $\gamma=\gamma(1)$ and the "alternating Euler constant"…

Classical Analysis and ODEs · Mathematics 2007-06-13 Jonathan Sondow , Petros Hadjicostas

This paper develops a general approach to characterize the long-time trajectory behavior of nonconvex gradient descent in generalized single-index models in the large aspect ratio regime. In this regime, we show that for each iteration the…

Machine Learning · Computer Science 2025-09-16 Qiyang Han

First we recall a method of computing scalar products of eigenfunctions of a Sturm-Liouville operator. This method is then applied to Macdonald and Gegenbauer functions, which are eigenfunctions of the Bessel, resp. Gegenbauer operators.…

Mathematical Physics · Physics 2024-05-17 Jan Dereziński , Christian Gaß , Błażej Ruba