English
Related papers

Related papers: Invariance Properties of the Natural Gradient in O…

200 papers

We study two procedures (reverse-mode and forward-mode) for computing the gradient of the validation error with respect to the hyperparameters of any iterative learning algorithm such as stochastic gradient descent. These procedures mirror…

Machine Learning · Statistics 2017-12-13 Luca Franceschi , Michele Donini , Paolo Frasconi , Massimiliano Pontil

Gradient methods are widely used in optimization problems. In practice, while the smoothness parameter can be estimated utilizing techniques such as backtracking, estimating the strong convexity parameter remains a challenge; moreover, even…

Optimization and Control · Mathematics 2026-02-17 Xiaozhe Hu , Sara Pollock , Zhongqin Xue , Yunrong Zhu

Multivariate functions encountered in high-dimensional uncertainty quantification problems often vary most strongly along a few dominant directions in the input parameter space. We propose a gradient-based method for detecting these…

Analysis of PDEs · Mathematics 2019-11-11 Olivier Zahm , Paul Constantine , Clémentine Prieur , Youssef Marzouk

Estimation of parameters in differential equation models can be achieved by applying learning algorithms to quantitative time-series data. However, sometimes it is only possible to measure qualitative changes of a system in response to a…

Machine Learning · Computer Science 2021-10-28 Gregory Szep , Neil Dalchau , Attila Csikasz-Nagy

Tuning hyperparameters of learning algorithms is hard because gradients are usually unavailable. We compute exact gradients of cross-validation performance with respect to all hyperparameters by chaining derivatives backwards through the…

Machine Learning · Statistics 2015-04-03 Dougal Maclaurin , David Duvenaud , Ryan P. Adams

Phase transitions with spontaneous symmetry breaking and vector order parameter are considered in multidimensional theory of general relativity. Covariant equations, describing the gravitational properties of topological defects, are…

General Relativity and Quantum Cosmology · Physics 2010-12-06 Boris E. Meierovich

The gradient technique is a promising tool with theoretical foundations based on the fundamental properties of MHD turbulence and turbulent reconnection. Its various incarnations use spectroscopic, synchrotron, and intensity data to trace…

Astrophysics of Galaxies · Physics 2024-06-04 Alex Lazarian , Ka Ho Yuen , Dmitri Pogosyan

Phase transitions with spontaneous symmetry breaking and vector order parameter are considered in multidimensional theory of general relativity. Covariant equations, describing the gravitational properties of topological defects, are…

General Relativity and Quantum Cosmology · Physics 2014-11-21 Boris E. Meierovich

This paper is devoted to the investigation of gradient flows in asymmetric metric spaces (for example, irreversible Finsler manifolds and Minkowski normed spaces) by means of discrete approximation. We study basic properties of curves and…

Differential Geometry · Mathematics 2023-07-21 Shin-ichi Ohta , Wei Zhao

Variational Bayes (VB) is a critical method in machine learning and statistics, underpinning the recent success of Bayesian deep learning. The natural gradient is an essential component of efficient VB estimation, but it is prohibitively…

Quantum Physics · Physics 2022-06-22 Anna Lopatnikova , Minh-Ngoc Tran

Adaptive gradient methods like AdaGrad are widely used in optimizing neural networks. Yet, existing convergence guarantees for adaptive gradient methods require either convexity or smoothness, and, in the smooth setting, only guarantee…

Machine Learning · Computer Science 2019-10-22 Xiaoxia Wu , Simon S. Du , Rachel Ward

Overparameterized shallow neural networks admit substantial parameter redundancy: distinct parameter vectors may represent the same predictor due to hidden-unit permutations, rescalings, and related symmetries. As a result, geometric…

Machine Learning · Computer Science 2026-03-24 Hang-Cheng Dong , Pengcheng Cheng

This paper shows that a wide class of effective learning rules -- those that improve a scalar performance measure over a given time window -- can be rewritten as natural gradient descent with respect to a suitably defined loss function and…

Machine Learning · Computer Science 2024-09-26 Lucas Shoji , Kenta Suzuki , Leo Kozachkov

We study the Wasserstein natural gradient in parametric statistical models with continuous sample spaces. Our approach is to pull back the $L^2$-Wasserstein metric tensor in the probability density space to a parameter space, equipping the…

Optimization and Control · Mathematics 2024-08-20 Yifan Chen , Wuchen Li

We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a…

Machine Learning · Computer Science 2026-04-07 El Mehdi Achour , Kathlén Kohn , Holger Rauhut

We define a number of natural (from geometric and combinatorial points of view) deformation spaces of valuations on finite graphs, and study functions over these deformation spaces. These functions include both direct metric invariants…

Combinatorics · Mathematics 2007-05-23 Dmitry Jakobson , Igor Rivin

Natural-gradient methods enable fast and simple algorithms for variational inference, but due to computational difficulties, their use is mostly limited to \emph{minimal} exponential-family (EF) approximations. In this paper, we extend…

Machine Learning · Statistics 2020-11-09 Wu Lin , Mohammad Emtiyaz Khan , Mark Schmidt

An open question in the Deep Learning community is why neural networks trained with Gradient Descent generalize well on real datasets even though they are capable of fitting random data. We propose an approach to answering this question…

Machine Learning · Computer Science 2020-02-26 Satrajit Chatterjee

Natural gradient methods have been used to optimise the parameters of probability distributions in a variety of settings, often resulting in fast-converging procedures. Unfortunately, for many distributions of interest, computing the…

Machine Learning · Statistics 2024-05-28 Jonathan So , Richard E. Turner

In this paper, we propose a geometric framework to analyze the convergence properties of gradient descent trajectories in the context of linear neural networks. We translate a well-known empirical observation of linear neural nets into a…

Machine Learning · Computer Science 2023-08-02 Yacine Chitour , Zhenyu Liao , Romain Couillet