English
Related papers

Related papers: Normalized Gradients for All

200 papers

We use Levy processes to generate joint prior distributions, and therefore penalty functions, for a location parameter as p grows large. This generalizes the class of local-global shrinkage rules based on scale mixtures of normals,…

Methodology · Statistics 2011-04-26 Nicholas G. Polson , James G. Scott

The conditions of relative smoothness and relative strong convexity were recently introduced for the analysis of Bregman gradient methods for convex optimization. We introduce a generalized left-preconditioning method for gradient descent,…

Optimization and Control · Mathematics 2020-12-09 Chris J. Maddison , Daniel Paulin , Yee Whye Teh , Arnaud Doucet

Minimization of a smooth function on a sphere or, more generally, on a smooth manifold, is the simplest non-convex optimization problem. It has a lot of applications. Our goal is to propose a version of the gradient projection algorithm for…

Optimization and Control · Mathematics 2019-06-28 Maxim Balashov , Boris Polyak , Andrey Tremba

Due to the non-smoothness of optimization problems in Machine Learning, generalized smoothness assumptions have been gaining a lot of attention in recent years. One of the most popular assumptions of this type is $(L_0,L_1)$-smoothness…

Optimization and Control · Mathematics 2024-12-30 Eduard Gorbunov , Nazarii Tupitsa , Sayantan Choudhury , Alen Aliev , Peter Richtárik , Samuel Horváth , Martin Takáč

Gradient boundedness up to the boundary for solutions to Dirichlet and Neumann problems for elliptic systems with Uhlenbeck type structure is established. Nonlinearities of possibly non-polynomial type are allowed, and minimal regularity on…

Analysis of PDEs · Mathematics 2012-12-27 Andrea Cianchi , Vladimir Maz'ya

We present a new kind of normalization theorem: linearization theorem for skew products. The normal form is a skew product again, with the fiber maps linear. It appears, that even in the smooth case, the conjugacy is only H\"older…

Dynamical Systems · Mathematics 2015-08-28 Yulij Ilyashenko , Olga Romaskevich

We develop an optimization algorithm suitable for Bayesian learning in complex models. Our approach relies on natural gradient updates within a general black-box framework for efficient training with limited model-specific derivations. It…

Machine Learning · Statistics 2022-12-13 Martin Magris , Mostafa Shabani , Alexandros Iosifidis

We consider gradient-based optimisation of wide, shallow neural networks, where the output of each hidden node is scaled by a positive parameter. The scaling parameters are non-identical, differing from the classical Neural Tangent Kernel…

Machine Learning · Statistics 2025-02-19 Francois Caron , Fadhel Ayed , Paul Jung , Hoil Lee , Juho Lee , Hongseok Yang

We introduce a new method to prove lower estimates for the approximation error of general linear operators with smooth range in terms of classical moduli of smoothness and related $K$-functionals. In addition, we explicitly show how to…

Classical Analysis and ODEs · Mathematics 2017-06-05 Johannes Nagler

The paper considers the problem of network-based computation of global minima in smooth nonconvex optimization problems. It is known that distributed gradient-descent-type algorithms can achieve convergence to the set of global minima by…

Optimization and Control · Mathematics 2019-10-24 Brian Swenson , Anirudh Sridhar , H. Vincent Poor

We present a new family of min-max optimization algorithms that automatically exploit the geometry of the gradient data observed at earlier iterations to perform more informative extra-gradient steps in later ones. Thanks to this adaptation…

Optimization and Control · Mathematics 2020-11-20 Kimon Antonakopoulos , E. Veronica Belmega , Panayotis Mertikopoulos

We introduce a new type of boundary conditions, {\it smooth boundary conditions}, for numerical studies of quantum lattice systems. In a number of circumstances, these boundary conditions have substantially smaller finite-size effects than…

Condensed Matter · Physics 2009-10-22 M. Vekic , S. R. White

In recent work, we introduced topological notions of simple normal crossings symplectic divisor and variety, showed that they are equivalent, in a suitable sense, to the corresponding geometric notions, and established a topological…

Symplectic Geometry · Mathematics 2019-08-27 Mohammad Farajzadeh Tehrani , Mark McLean , Aleksey Zinger

We consider the generalization error associated with stochastic gradient descent on a smooth convex function over a compact set. We show the first bound on the generalization error that vanishes when the number of iterations $T$ and the…

Machine Learning · Computer Science 2024-04-16 Julien Hendrickx , Alex Olshevsky

Many theoretical results in deep learning can be traced to symmetry or equivariance of neural networks under parameter transformations. However, existing analyses are typically problem-specific and focus on first-order consequences such as…

Machine Learning · Computer Science 2025-12-29 Yongyi Yang , Liu Ziyin

We propose an algorithm to estimate the path-gradient of both the reverse and forward Kullback-Leibler divergence for an arbitrary manifestly invertible normalizing flow. The resulting path-gradient estimators are straightforward to…

Machine Learning · Computer Science 2022-07-19 Lorenz Vaitl , Kim A. Nicoli , Shinichi Nakajima , Pan Kessel

This paper studies first-order algorithms for solving fully composite optimization problems over convex and compact sets. We leverage the structure of the objective by handling its differentiable and non-differentiable components…

Optimization and Control · Mathematics 2023-07-13 Maria-Luiza Vladarean , Nikita Doikov , Martin Jaggi , Nicolas Flammarion

We are concerned with local regularity of the solutions for the Stokes and Navier-Stokes equations near boundary. Firstly, we construct a bounded solution but its normal derivatives are singular in any $L^p$ with $1<p$ locally near…

Analysis of PDEs · Mathematics 2022-04-18 Tongkeun Chang , Kyungkeun Kang

Generalization error (also known as the out-of-sample error) measures how well the hypothesis learned from training data generalizes to previously unseen data. Proving tight generalization error bounds is a central question in statistical…

Machine Learning · Computer Science 2020-03-03 Jian Li , Xuanyuan Luo , Mingda Qiao

For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization. Unfortunately, for deep models, this relationship…

Machine Learning · Computer Science 2021-06-17 Colin Wei , Tengyu Ma