English
Related papers

Related papers: Invariance Properties of the Natural Gradient in O…

200 papers

We explicitly construct parameter transformations between gradient flows in metric spaces, called curves of maximal slope, having different exponents when the associated function satisfies a suitable convexity condition. These…

Analysis of PDEs · Mathematics 2024-04-04 Sho Shimoyama

Fine-tuning and naturalness, the sensitivity of low-energy observables to small changes in the fundamental parameters of a theory, are cornerstones of physics beyond the Standard Model. We propose a new measure of fine-tuning based on…

High Energy Physics - Theory · Physics 2026-05-04 James Halverson , Thomas R. Harvey , Michael Nee

Stochastic models with global parameters and latent variables are common, and for which variational inference (VI) is popular. However, existing methods are often either slow or inaccurate in high dimensions. We suggest a fast and accurate…

Machine Learning · Statistics 2024-07-26 Weiben Zhang , Michael Stanley Smith , Worapree Maneesoonthorn , Ruben Loaiza-Maya

The gauge-invariant two-point function of the Higgs field at the same spacetime point can make a natural gauge-invariant order parameter for spontaneous gauge symmetry breaking. However, this composite operator is ultraviolet divergent and…

High Energy Physics - Theory · Physics 2023-06-14 Kengo Kikuchi , Kenji Nishiwaki , Kin-ya Oda

Matrix-valued covariance functions are crucial to geostatistical modeling of multivariate spatial data. The classical assumption of symmetry of a multivariate covariance function is overlay restrictive and has been considered as unrealistic…

Statistics Theory · Mathematics 2017-11-28 Alfredo Alegría , Emilio Porcu , Reinhard Furrer

Natural gradients have been widely used in optimization of loss functionals over probability space, with important examples such as Fisher-Rao gradient descent for Kullback-Leibler divergence, Wasserstein gradient descent for…

Numerical Analysis · Mathematics 2020-06-30 Lexing Ying

A quantum generalization of Natural Gradient Descent is presented as part of a general-purpose optimization framework for variational quantum circuits. The optimization dynamics is interpreted as moving in the steepest descent direction…

Quantum Physics · Physics 2020-05-27 James Stokes , Josh Izaac , Nathan Killoran , Giuseppe Carleo

This paper proposes a gradient descent based optimization method that relies on automatic differentiation for the computation of gradients. The method uses tools and techniques originally developed in the field of artificial neural networks…

Systems and Control · Electrical Eng. & Systems 2023-09-29 Georg Kordowich , Johann Jaeger

The remarkable practical success of deep learning has revealed some major surprises from a theoretical perspective. In particular, simple gradient methods easily find near-optimal solutions to non-convex optimization problems, and despite…

Statistics Theory · Mathematics 2021-03-17 Peter L. Bartlett , Andrea Montanari , Alexander Rakhlin

Kakade's natural policy gradient method has been studied extensively in recent years, showing linear convergence with and without regularization. We study another natural gradient method based on the Fisher information matrix of the…

Optimization and Control · Mathematics 2025-02-05 Johannes Müller , Semih Çaycı , Guido Montúfar

Many modern learning tasks involve fitting nonlinear models to data which are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset. Due to this overparameterization, the training…

Machine Learning · Computer Science 2018-12-27 Samet Oymak , Mahdi Soltanolkotabi

We consider machine learning tasks with low-rank functional tree tensor networks (TTN) as the learning model. While in the case of least-squares regression, low-rank functional TTNs can be efficiently optimized using alternating…

Optimization and Control · Mathematics 2026-04-13 Nikolas Klug , Michael Ulbrich , André Uschmajew , Marius Willner

Thermal states play a fundamental role in various areas of physics, and they are becoming increasingly important in quantum information science, with applications related to semi-definite programming, quantum Boltzmann machine learning,…

Quantum Physics · Physics 2025-11-13 Dhrumil Patel , Mark M. Wilde

Feedback alignment algorithms are an alternative to backpropagation to train neural networks, whereby some of the partial derivatives that are required to compute the gradient are replaced by random terms. This essentially transforms the…

Machine Learning · Computer Science 2023-06-06 Dominique Chu , Florian Bacho

We study diffusions, variational principles and associated boundary value problems on directed graphs with natural weightings. Using random walks and exit times, we associate to certain subgraphs (domains) a pair of sequences, each of which…

Spectral Theory · Mathematics 2007-05-23 Patrick McDonald , Robert Meyers

Implicit regularization refers to the tendency of local search algorithms to converge to low-dimensional solutions, even when such structures are not explicitly enforced. Despite its ubiquity, the mechanism underlying this behavior remains…

Machine Learning · Computer Science 2025-12-10 Jianhao Ma , Geyu Liang , Salar Fattahi

Natural policy gradient (NPG) and its variants are widely-used policy search methods in reinforcement learning. Inspired by prior work, a new NPG variant coined NPG-HM is developed in this paper, which utilizes the Hessian-aided momentum…

Machine Learning · Computer Science 2024-01-23 Jie Feng , Ke Wei , Jinchi Chen

We describe four algorithms for neural network training, each adapted to different scalability constraints. These algorithms are mathematically principled and invariant under a number of transformations in data and network representation,…

Neural and Evolutionary Computing · Computer Science 2015-02-04 Yann Ollivier

Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with parameter averaging,…

Machine Learning · Statistics 2018-11-22 Dmitry Babichev , Francis Bach

We study the gradient method under the assumption that an additively inexact gradient is available for, generally speaking, non-convex problems. The non-convexity of the objective function, as well as the use of an inexactness specified…

Optimization and Control · Mathematics 2022-12-13 Boris T. Polyak , Ilia A. Kuruzov , Fedor S. Stonyakin