Related papers: Normalized Gradients for All
We use Levy processes to generate joint prior distributions, and therefore penalty functions, for a location parameter as p grows large. This generalizes the class of local-global shrinkage rules based on scale mixtures of normals,…
The conditions of relative smoothness and relative strong convexity were recently introduced for the analysis of Bregman gradient methods for convex optimization. We introduce a generalized left-preconditioning method for gradient descent,…
Minimization of a smooth function on a sphere or, more generally, on a smooth manifold, is the simplest non-convex optimization problem. It has a lot of applications. Our goal is to propose a version of the gradient projection algorithm for…
Due to the non-smoothness of optimization problems in Machine Learning, generalized smoothness assumptions have been gaining a lot of attention in recent years. One of the most popular assumptions of this type is $(L_0,L_1)$-smoothness…
Gradient boundedness up to the boundary for solutions to Dirichlet and Neumann problems for elliptic systems with Uhlenbeck type structure is established. Nonlinearities of possibly non-polynomial type are allowed, and minimal regularity on…
We present a new kind of normalization theorem: linearization theorem for skew products. The normal form is a skew product again, with the fiber maps linear. It appears, that even in the smooth case, the conjugacy is only H\"older…
We develop an optimization algorithm suitable for Bayesian learning in complex models. Our approach relies on natural gradient updates within a general black-box framework for efficient training with limited model-specific derivations. It…
We consider gradient-based optimisation of wide, shallow neural networks, where the output of each hidden node is scaled by a positive parameter. The scaling parameters are non-identical, differing from the classical Neural Tangent Kernel…
We introduce a new method to prove lower estimates for the approximation error of general linear operators with smooth range in terms of classical moduli of smoothness and related $K$-functionals. In addition, we explicitly show how to…
The paper considers the problem of network-based computation of global minima in smooth nonconvex optimization problems. It is known that distributed gradient-descent-type algorithms can achieve convergence to the set of global minima by…
We present a new family of min-max optimization algorithms that automatically exploit the geometry of the gradient data observed at earlier iterations to perform more informative extra-gradient steps in later ones. Thanks to this adaptation…
We introduce a new type of boundary conditions, {\it smooth boundary conditions}, for numerical studies of quantum lattice systems. In a number of circumstances, these boundary conditions have substantially smaller finite-size effects than…
In recent work, we introduced topological notions of simple normal crossings symplectic divisor and variety, showed that they are equivalent, in a suitable sense, to the corresponding geometric notions, and established a topological…
We consider the generalization error associated with stochastic gradient descent on a smooth convex function over a compact set. We show the first bound on the generalization error that vanishes when the number of iterations $T$ and the…
Many theoretical results in deep learning can be traced to symmetry or equivariance of neural networks under parameter transformations. However, existing analyses are typically problem-specific and focus on first-order consequences such as…
We propose an algorithm to estimate the path-gradient of both the reverse and forward Kullback-Leibler divergence for an arbitrary manifestly invertible normalizing flow. The resulting path-gradient estimators are straightforward to…
This paper studies first-order algorithms for solving fully composite optimization problems over convex and compact sets. We leverage the structure of the objective by handling its differentiable and non-differentiable components…
We are concerned with local regularity of the solutions for the Stokes and Navier-Stokes equations near boundary. Firstly, we construct a bounded solution but its normal derivatives are singular in any $L^p$ with $1<p$ locally near…
Generalization error (also known as the out-of-sample error) measures how well the hypothesis learned from training data generalizes to previously unseen data. Proving tight generalization error bounds is a central question in statistical…
For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization. Unfortunately, for deep models, this relationship…