Related papers: Normalized Gradients for All
Machine learning models trained by different optimization algorithms under different data distributions can exhibit distinct generalization behaviors. In this paper, we analyze the generalization of models trained by noisy iterative…
Hyperbolic neural networks (HNNs) have demonstrated notable efficacy in representing real-world data with hierarchical structures via exploiting the geometric properties of hyperbolic spaces characterized by negative curvatures. Curvature…
For a locally Lipschitz continuous function $f:X\to\mathbb{R}$ the generalized gradient $\partial f(x)$ of Clarke is used to develop some (set-valued) gradient on a set $A\subset X$. Existence, uniqueness and some approximation are…
Adaptive optimization methods have been widely used in deep learning. They scale the learning rates adaptively according to the past gradient, which has been shown to be effective to accelerate the convergence. However, they suffer from…
Relaxation of initially out-of-equilibrium rough interfaces in presence of thermal noise is investigated using Langevin formalism. During thermal equilibration towards the well-known roughening regime, three scaling regimes observed over…
This paper investigates the relation between the boundary geometric properties and the boundary regularity of the solutions of elliptic equations. We prove by a new unified method the pointwise boundary H\"{o}lder regularity under proper…
We formulate and prove $\textit{a priori}$ bounds for the renormalization of H\'enon-like maps (under certain regularity assumptions). This provides a certain uniform control on the small-scale geometry of the dynamics, and ensures…
Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving…
Recently there were proposed some innovative convex optimization concepts, namely, relative smoothness [1] and relative strong convexity [2,3]. These approaches have significantly expanded the class of applicability of gradient-type methods…
A theory of graded manifolds can be viewed as a generalization of differential geometry of smooth manifolds. It allows one to work with functions which locally depend not only on ordinary real variables, but also on $\mathbb{Z}$-graded…
An algorithm is said to be adaptive to a certain parameter (of the problem) if it does not need a priori knowledge of such a parameter but performs competitively to those that know it. This dissertation presents our work on adaptive…
We generalize the notion of average Lipschitz smoothness proposed by Ashlagi et al. (COLT 2021) by extending it to H\"older smoothness. This measure of the "effective smoothness" of a function is sensitive to the underlying distribution and…
Variational approaches to disparity estimation typically use a linearised brightness constancy constraint, which only applies in smooth regions and over small distances. Accordingly, current variational approaches rely on a schedule to…
Black-box variational inference is widely used in situations where there is no proof that its stochastic optimization succeeds. We suggest this is due to a theoretical gap in existing stochastic optimization proofs: namely the challenge of…
We propose to use the {\L}ojasiewicz inequality as a general tool for analyzing the convergence rate of gradient descent on a Hilbert manifold, without resorting to the continuous gradient flow. Using this tool, we show that a Sobolev…
In this article, by applying the well known method for dealing with $p$-Laplace type elliptic boundary value problems, the authors establish a sharp estimate for the decreasing rearrangement of the gradient of solutions to the Dirichlet and…
We show that for any uniformly elliptic fully nonlinear second-order equation with bounded measurable "coefficients" and bounded "free" term one can find an approximating equation which has a unique continuous and having the second…
Gradient-based minimax optimal algorithms have greatly promoted the development of continuous optimization and machine learning. One seminal work due to Yurii Nesterov [Nes83a] established $\tilde{\mathcal{O}}(\sqrt{L/\mu})$ gradient…
This work develops a mean-field analysis for the asymptotic behavior of deep BitNet-like architectures as smooth quantization parameters approach zero. We establish that empirical measures of latent weights converge weakly to solutions of…
We treat the heat equation with singular drift terms and its generalization: the linearized Navier-Stokes system. In the first case, we obtain boundedness of weak solutions for highly singular, "supercritical" data. In the second case, we…