English
Related papers

Related papers: Invariance Properties of the Natural Gradient in O…

200 papers

We consider the dynamics of gradient descent (GD) in overparameterized single hidden layer neural networks with a squared loss function. Recently, it has been shown that, under some conditions, the parameter values obtained using GD achieve…

Machine Learning · Computer Science 2021-05-17 Siddhartha Satpathi , R Srikant

Multivariate spatial field data are increasingly common and whose modeling typically relies on building cross-covariance functions to describe cross-process relationships. An alternative viewpoint is to model the matrix of spectral…

Statistics Theory · Mathematics 2015-05-07 William Kleiber

In order to develop a differential calculus for error propagation we study local Dirichlet forms on probability spaces with square field operator $\Gamma$ -- i.e. error structures -- and we are looking for an object related to $\Gamma$…

Probability · Mathematics 2007-05-23 Nicolas Bouleau

Adaptive gradient methods have achieved remarkable success in training deep neural networks on a wide variety of tasks. However, not much is known about the mathematical and statistical properties of this family of methods. This work aims…

Machine Learning · Computer Science 2021-05-18 Zhang Zhiyi , Liu Ziyin

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large…

Machine Learning · Computer Science 2025-10-21 Lachlan Ewen MacDonald , Hancheng Min , Leandro Palma , Salma Tarmoun , Ziqing Xu , René Vidal

This paper explores variants of the subspace iteration algorithm for computing approximate invariant subspaces. The standard subspace iteration approach is revisited and new variants that exploit gradient-type techniques combined with a…

Numerical Analysis · Mathematics 2024-05-14 Foivos Alimisis , Yousef Saad , Bart Vandereycken

We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or…

Machine Learning · Statistics 2020-06-24 Suriya Gunasekar , Jason Lee , Daniel Soudry , Nathan Srebro

It is important in many applications to be able to extend the (outer) unit normal vector field from a hypersurface to its neighborhood in such a way that the result is a unit gradient field. The aim of the paper is to provide an elementary…

Differential Geometry · Mathematics 2018-02-16 R. Duduchava , E. Shargorodsky , G. Tephnadze

We propose a new technique that boosts the convergence of training generative adversarial networks. Generally, the rate of training deep models reduces severely after multiple iterations. A key reason for this phenomenon is that a deep…

Machine Learning · Statistics 2018-06-15 Atsushi Nitanda , Taiji Suzuki

In many relevant cases -- e.g., in hamiltonian dynamics -- a given vector field can be characterized by means of a variational principle based on a one-form. We discuss how a vector field on a manifold can also be characterized in a similar…

Mathematical Physics · Physics 2015-06-26 G. Gaeta , P. Morando

As the demand to integrate Artificial Intelligence into high-stakes environments continues to grow, explaining the reasoning behind neural-network predictions has shifted from a theoretical curiosity to a strict operational requirement. Our…

Machine Learning · Statistics 2026-04-27 Younes Essafouri , Laure Raynaud , Luciano Drozda , Laurent Risser

While the optimization landscape of policy gradient methods has been recently investigated for partially observed linear systems in terms of both static output feedback and dynamical controllers, they only provide convergence guarantees to…

Optimization and Control · Mathematics 2023-04-25 Feiran Zhao , Xingyun Fu , Keyou You

The paper proposes a variational-inequality based primal-dual dynamic that has a globally exponentially stable saddle-point solution when applied to solve linear inequality constrained optimization problems. A Riemannian geometric framework…

Optimization and Control · Mathematics 2020-10-07 P. Bansode , V. Chinde , S. R. Wagh , R. Pasumarthy , N. M. Singh

We study the convergence of several natural policy gradient (NPG) methods in infinite-horizon discounted Markov decision processes with regular policy parametrizations. For a variety of NPGs and reward functions we show that the…

Optimization and Control · Mathematics 2024-02-21 Johannes Müller , Guido Montúfar

We consider the effect of a random longitudinal field on the Ising model in a transverse magnetic field. For spatial dimension $d > 2$, there is at low strength of randomness and transverse field, a phase with true long range order which is…

Disordered Systems and Neural Networks · Physics 2016-08-31 T. Senthil

This paper presents enhancement strategies for the Hermitian and skew-Hermitian splitting method based on gradient iterations. The spectral properties are exploited for the parameter estimation, often resulting in a better convergence. In…

Numerical Analysis · Mathematics 2020-07-08 Qinmeng Zou , Frederic Magoules

We consider a general formulation of gradient flow evolution for problems whose natural framework is the one of metric spaces. The applications we deal with are concerned with the evolution of {\it capacitary measures} with respect to the…

Analysis of PDEs · Mathematics 2011-09-27 Dorin Bucur , Giuseppe Buttazzo , Ulisse Stefanelli

What is the optimal way to approximate a high-dimensional diffusion process by one in which the coordinates are independent? This paper presents a construction, called the \emph{independent projection}, which is optimal for two natural…

Probability · Mathematics 2024-10-10 Daniel Lacker

This paper introduces a method for efficiently approximating the inverse of the Fisher information matrix, a crucial step in achieving effective variational Bayes inference. A notable aspect of our approach is the avoidance of analytically…

Methodology · Statistics 2024-04-29 A. Godichon-Baggioni , D. Nguyen , M-N Tran

We provide a theoretical explanation for the effectiveness of gradient clipping in training deep neural networks. The key ingredient is a new smoothness condition derived from practical neural network training examples. We observe that…

Optimization and Control · Mathematics 2020-02-12 Jingzhao Zhang , Tianxing He , Suvrit Sra , Ali Jadbabaie