English
Related papers

Related papers: Normalized Gradients for All

200 papers

We provide the first proof of convergence for normalized error feedback algorithms across a wide range of machine learning problems. Despite their popularity and efficiency in training deep neural networks, traditional analyses of error…

Machine Learning · Computer Science 2024-10-23 Sarit Khirirat , Abdurakhmon Sadiev , Artem Riabinin , Eduard Gorbunov , Peter Richtárik

We prove an analogue for a one-phase free boundary problem of the classical gradient bound for solutions to the minimal surface equation. It follows, in particular, that every energy-minimizing free boundary that is a graph is also smooth.…

Analysis of PDEs · Mathematics 2010-09-24 Daniela De Silva , David Jerison

Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds…

Machine Learning · Statistics 2026-05-11 Mana Sakai , Masaaki Imaizumi

We evaluate natural gradient, an algorithm originally proposed in Amari (1997), for learning deep models. The contributions of this paper are as follows. We show the connection between natural gradient and three other recently proposed…

Machine Learning · Computer Science 2014-02-18 Razvan Pascanu , Yoshua Bengio

In this paper, we prove new complexity bounds for zeroth-order methods in non-convex optimization with inexact observations of the objective function values. We use the Gaussian smoothing approach of Nesterov and Spokoiny [2015] and extend…

Optimization and Control · Mathematics 2021-01-14 Innokentiy Shibaev , Pavel Dvurechensky , Alexander Gasnikov

We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature. No need for functional values, no line search, no…

Optimization and Control · Mathematics 2020-08-18 Yura Malitsky , Konstantin Mishchenko

Variational methods for revealing visual concepts learned by convolutional neural networks have gained significant attention during the last years. Being based on noisy gradients obtained via back-propagation such methods require the…

Machine Learning · Computer Science 2018-05-02 Maximilian Baust , Florian Ludwig , Christian Rupprecht , Matthias Kohl , Stefan Braunewell

In this paper, we deal with the problem of optimizing a black-box smooth function over a full-dimensional smooth convex set. We study sets of feasible curves that allow to properly characterize stationarity of a solution and possibly carry…

Optimization and Control · Mathematics 2026-04-03 Xiaoxi Jia , Matteo Lapucci , Pierluigi Mansueto

Motivated by a wide variety of applications, ranging from stochastic optimization to dimension reduction through variable selection, the problem of estimating gradients accurately is of crucial importance in statistics and learning theory.…

Machine Learning · Computer Science 2020-06-29 Guillaume Ausset , Stephan Clémençon , François Portier

In this paper, we consider linear elliptic systems from composite materials where the coefficients depend on the shape and might have the discontinuity between the subregions. We derive a function which is related to the gradient of the…

Analysis of PDEs · Mathematics 2022-06-17 Youchan Kim , Pilsoo Shin

We provide improved error bounds for kernel-based numerical differentiation in terms of growth functions when kernels are of a finite smoothness, such as polyharmonic splines, thin plate splines or Wendland kernels. In contrast to existing…

Numerical Analysis · Mathematics 2025-12-24 Oleg Davydov

We derive a new generalization of the nonlinear variational wave equation. We prove existence of local, smooth solutions for this system. As a limiting case, we recover the nonlinear variational wave equation.

Analysis of PDEs · Mathematics 2023-08-15 Katrin Grunert , Audun Reigstad

Integrated gradients is prevalent within machine learning to address the black-box problem of neural networks. The explanations given by integrated gradients depend on a choice of base-point. The choice of base-point is not a priori obvious…

Machine Learning · Computer Science 2025-03-12 Lachlan Simpson , Federico Costanza , Kyle Millar , Adriel Cheng , Cheng-Chew Lim , Hong Gunn Chew

Universal methods for optimization are designed to achieve theoretically optimal convergence rates without any prior knowledge of the problem's regularity parameters or the accurarcy of the gradient oracle employed by the optimizer. In this…

Optimization and Control · Mathematics 2022-06-22 Kimon Antonakopoulos , Dong Quan Vu , Vokan Cevher , Kfir Y. Levy , Panayotis Mertikopoulos

We study the higher H\"older regularity of local weak solutions to a class of nonlinear nonlocal elliptic equations with kernels that satisfy a mild continuity assumption. An interesting feature of our main result is that the obtained…

Analysis of PDEs · Mathematics 2021-01-19 Simon Nowak

The goal of this paper is to debunk and dispel the magic behind black-box optimizers and stochastic optimizers. It aims to build a solid foundation on how and why the techniques work. This manuscript crystallizes this knowledge by deriving…

Machine Learning · Computer Science 2024-01-15 Jun Lu

We introduce an algorithm to solve linear inverse problems regularized with the total (gradient) variation in a gridless manner. Contrary to most existing methods, that produce an approximate solution which is piecewise constant on a fixed…

Signal Processing · Electrical Eng. & Systems 2025-07-08 Yohann de Castro , Vincent Duval , Romain Petit

In this paper, we investigate interior gradient estimates for solutions to the mean curvature equation $$ \dive \left( \frac{\nabla u}{\sqrt{1 + |\nabla u|^2}} \right) = f(\nabla u)$$ under various nonlinear assumptions on the right-hand…

Analysis of PDEs · Mathematics 2026-02-13 Fanheng Xu

The key to generalization is controlling the complexity of the network. However, there is no obvious control of complexity -- such as an explicit regularization term -- in the training of deep networks for classification. We will show that…

Machine Learning · Computer Science 2020-04-14 Andrzej Banburski , Qianli Liao , Brando Miranda , Lorenzo Rosasco , Fernanda De La Torre , Jack Hidary , Tomaso Poggio

Zeroth-order (ZO) optimization is popular in real-world applications that accessing the gradient information is expensive or unavailable. Recently, adaptive ZO methods that normalize gradient estimators by the empirical standard deviation…

Optimization and Control · Mathematics 2026-02-03 Haishan Ye , Luo Luo