English
Related papers

Related papers: Error whitening: Why Gauss-Newton outperforms Newt…

200 papers

In this paper, we study random subsampling of Gaussian process regression, one of the simplest approximation baselines, from a theoretical perspective. Although subsampling discards a large part of training data, we show provable guarantees…

Machine Learning · Statistics 2019-01-29 Kohei Hayashi , Masaaki Imaizumi , Yuichi Yoshida

In this paper we model the loss function of high-dimensional optimization problems by a Gaussian random field, or equivalently a Gaussian process. Our aim is to study gradient descent in such loss functions or energy landscapes and compare…

Machine Learning · Statistics 2018-03-28 Mariano Chouza , Stephen Roberts , Stefan Zohren

In this paper we present GSSN, a globalized SCD semismooth* Newton method for solving nonsmooth nonconvex optimization problems. The global convergence properties of the method are ensured by the proximal gradient method, whereas locally…

Optimization and Control · Mathematics 2025-01-27 H. Gfrerer

Natural Gradient Descent, a second-degree optimization method motivated by the information geometry, makes use of the Fisher Information Matrix instead of the Hessian which is typically used. However, in many cases, the Fisher Information…

Machine Learning · Computer Science 2023-03-10 Rajesh Shrestha

In this paper, we present some theoretical work to explain why simple gradient descent methods are so successful in solving non-convex optimization problems in learning large-scale neural networks (NN). After introducing a mathematical tool…

Machine Learning · Computer Science 2023-05-01 Hui Jiang

We propose a novel class of neural network-like parametrized functions, i.e., general transformation neural networks (GTNNs), for high-dimensional approximation. Conventional deep neural networks sometimes perform less accurately on…

Numerical Analysis · Mathematics 2026-02-25 Xiaoyang Wang , Yiqi Gu

We consider a variant of matrix completion where entries are revealed in a biased manner. We wish to understand the extent to which such bias can be exploited in improving predictions. Towards that, we propose a natural model where the…

Machine Learning · Computer Science 2025-01-03 Yassir Jedra , Sean Mann , Charlotte Park , Devavrat Shah

This manuscript considers the problem of learning a random Gaussian network function using a fully connected network with frozen intermediate layers and trainable readout layer. This problem can be seen as a natural generalization of the…

Machine Learning · Statistics 2023-02-02 Dominik Schröder , Hugo Cui , Daniil Dmitriev , Bruno Loureiro

Gaussian processes are a fully Bayesian smoothing technique that allows for the reconstruction of a function and its derivatives directly from observational data, without assuming a specific model or choosing a parameterization. This is…

Cosmology and Nongalactic Astrophysics · Physics 2013-11-27 Marina Seikel , Chris Clarkson

In this paper we generalize the technique of deflation to define two new methods to systematically find many local minima of a nonlinear least squares problem. The methods are based on the Gauss-Newton algorithm, and as such do not require…

Numerical Analysis · Mathematics 2025-06-13 Alban Bloor Riley , Marcus Webb , Michael L Baker

We study a semismooth Newton-type method for the nearest doubly stochastic matrix problem where both differentiability and nonsingularity of the Jacobian can fail. The optimality conditions for this problem are formulated as a system of…

Optimization and Control · Mathematics 2021-07-21 Hao Hu , Haesol Im , Xinxin Li , Henry Wolkowicz

In this paper, we consider a modified projected Gauss-Newton method for solving constrained nonlinear least-squares problems. We assume that the functional constraints are smooth and the the other constraints are represented by a simple…

Optimization and Control · Mathematics 2025-04-02 Yassine Nabou , Lucian Toma , Ion Necoara

By learning the gradient of smoothed data distributions, diffusion models can iteratively generate samples from complex distributions. The learned score function enables their generalization capabilities, but how the learned score relates…

Machine Learning · Computer Science 2024-12-16 Binxu Wang , John J. Vastola

This paper presents a detailed discussion of the ``Newton's method'' algorithm for finding apparent horizons in 3+1 numerical relativity. We describe a method for computing the Jacobian matrix of the finite differenced $H(h)$ function by…

General Relativity and Quantum Cosmology · Physics 2009-07-10 Jonathan Thornburg

Robustness, domain adaptation, photometric/occlusion invariance, sensor drift, and alignment style are treated as separate literatures with separate method families. Under label-preserving deployment shift they share one geometric object:…

Machine Learning · Computer Science 2026-05-26 Vishal Rajput

It is generally known that counting statistics is not correctly described by a Gaussian approximation. Nevertheless, in neutron scattering, it is common practice to apply this approximation to the counting statistics; also at low counting…

Data Analysis, Statistics and Probability · Physics 2020-06-09 Jakob Lassa , Magnus Egede Bøggild , Per Hedegård , Kim Lefmann

For strongly convex objectives that are smooth, the classical theory of gradient descent ensures linear convergence relative to the number of gradient evaluations. An analogous nonsmooth theory is challenging. Even when the objective is…

Optimization and Control · Mathematics 2023-01-19 X. Y. Han , Adrian S. Lewis

We introduce the Gaussian transform (GT), an optimal transport inspired iterative method for denoising and enhancing latent structures in datasets. Under the hood, GT generates a new distance function (GT distance) on a given dataset by…

Machine Learning · Computer Science 2020-06-23 Kun Jin , Facundo Mémoli , Zhengchao Wan

Modern deep learning models have achieved great success in predictive accuracy for many data modalities. However, their application to many real-world tasks is restricted by poor uncertainty estimates, such as overconfidence on…

Machine Learning · Statistics 2020-10-16 Ben Adlam , Jaehoon Lee , Lechao Xiao , Jeffrey Pennington , Jasper Snoek

Predict and optimize is an increasingly popular decision-making paradigm that employs machine learning to predict unknown parameters of optimization problems. Instead of minimizing the prediction error of the parameters, it trains…

Machine Learning · Computer Science 2024-02-05 Grigorii Veviurko , Wendelin Böhmer , Mathijs de Weerdt