English
Related papers

Related papers: Normalized Gradients for All

200 papers

Machine learning models trained by different optimization algorithms under different data distributions can exhibit distinct generalization behaviors. In this paper, we analyze the generalization of models trained by noisy iterative…

Machine Learning · Statistics 2022-12-29 Hao Wang , Rui Gao , Flavio P. Calmon

Hyperbolic neural networks (HNNs) have demonstrated notable efficacy in representing real-world data with hierarchical structures via exploiting the geometric properties of hyperbolic spaces characterized by negative curvatures. Curvature…

Machine Learning · Computer Science 2025-08-27 Xiaomeng Fan , Yuwei Wu , Zhi Gao , Mehrtash Harandi , Yunde Jia

For a locally Lipschitz continuous function $f:X\to\mathbb{R}$ the generalized gradient $\partial f(x)$ of Clarke is used to develop some (set-valued) gradient on a set $A\subset X$. Existence, uniqueness and some approximation are…

Optimization and Control · Mathematics 2018-03-19 Jan Mankau , Friedemann Schuricht

Adaptive optimization methods have been widely used in deep learning. They scale the learning rates adaptively according to the past gradient, which has been shown to be effective to accelerate the convergence. However, they suffer from…

Machine Learning · Computer Science 2021-07-06 Hongwei Zhang , Weidong Zou , Hongbo Zhao , Qi Ming , Tijin Yan , Yuanqing Xia , Weipeng Cao

Relaxation of initially out-of-equilibrium rough interfaces in presence of thermal noise is investigated using Langevin formalism. During thermal equilibration towards the well-known roughening regime, three scaling regimes observed over…

Statistical Mechanics · Physics 2010-08-25 Thi Thu Thuy Nguyen , Daniel Bonamy , Laurent Phan Vam , Jacques Cousty , Luc Barbier

This paper investigates the relation between the boundary geometric properties and the boundary regularity of the solutions of elliptic equations. We prove by a new unified method the pointwise boundary H\"{o}lder regularity under proper…

Analysis of PDEs · Mathematics 2020-06-16 Yuanyuan Lian , Kai Zhang , Dongsheng Li , Guanghao Hong

We formulate and prove $\textit{a priori}$ bounds for the renormalization of H\'enon-like maps (under certain regularity assumptions). This provides a certain uniform control on the small-scale geometry of the dynamics, and ensures…

Dynamical Systems · Mathematics 2024-11-22 Sylvain Crovisier , Mikhail Lyubich , Enrique Pujals , Jonguk Yang

Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving…

Machine Learning · Computer Science 2018-05-25 Jun Li , Hongfu Liu , Bineng Zhong , Yue Wu , Yun Fu

Recently there were proposed some innovative convex optimization concepts, namely, relative smoothness [1] and relative strong convexity [2,3]. These approaches have significantly expanded the class of applicability of gradient-type methods…

Optimization and Control · Mathematics 2024-04-19 Fedor Stonyakin , Alexander Titov , Mohammad Alkousa , Oleg Savchuk , Alexander Gasnikov

A theory of graded manifolds can be viewed as a generalization of differential geometry of smooth manifolds. It allows one to work with functions which locally depend not only on ordinary real variables, but also on $\mathbb{Z}$-graded…

Differential Geometry · Mathematics 2023-03-14 Jan Vysoky

An algorithm is said to be adaptive to a certain parameter (of the problem) if it does not need a priori knowledge of such a parameter but performs competitively to those that know it. This dissertation presents our work on adaptive…

Machine Learning · Computer Science 2023-07-10 Zhenxun Zhuang

We generalize the notion of average Lipschitz smoothness proposed by Ashlagi et al. (COLT 2021) by extending it to H\"older smoothness. This measure of the "effective smoothness" of a function is sensitive to the underlying distribution and…

Machine Learning · Computer Science 2023-10-31 Steve Hanneke , Aryeh Kontorovich , Guy Kornowski

Variational approaches to disparity estimation typically use a linearised brightness constancy constraint, which only applies in smooth regions and over small distances. Accordingly, current variational approaches rely on a schedule to…

Image and Video Processing · Electrical Eng. & Systems 2024-05-28 James L. Gray , Aous T. Naman , David S. Taubman

Black-box variational inference is widely used in situations where there is no proof that its stochastic optimization succeeds. We suggest this is due to a theoretical gap in existing stochastic optimization proofs: namely the challenge of…

Machine Learning · Computer Science 2023-12-25 Justin Domke , Guillaume Garrigos , Robert Gower

We propose to use the {\L}ojasiewicz inequality as a general tool for analyzing the convergence rate of gradient descent on a Hilbert manifold, without resorting to the continuous gradient flow. Using this tool, we show that a Sobolev…

Numerical Analysis · Mathematics 2021-05-21 Ziyun Zhang

In this article, by applying the well known method for dealing with $p$-Laplace type elliptic boundary value problems, the authors establish a sharp estimate for the decreasing rearrangement of the gradient of solutions to the Dirichlet and…

Analysis of PDEs · Mathematics 2016-03-03 Sibei Yang , Der-Chen Chang , Dachun Yang , Zunwei Fu

We show that for any uniformly elliptic fully nonlinear second-order equation with bounded measurable "coefficients" and bounded "free" term one can find an approximating equation which has a unique continuous and having the second…

Analysis of PDEs · Mathematics 2012-04-03 N. V. Krylov

Gradient-based minimax optimal algorithms have greatly promoted the development of continuous optimization and machine learning. One seminal work due to Yurii Nesterov [Nes83a] established $\tilde{\mathcal{O}}(\sqrt{L/\mu})$ gradient…

Machine Learning · Computer Science 2023-12-07 Yuanshi Liu , Hanzhen Zhao , Yang Xu , Pengyun Yue , Cong Fang

This work develops a mean-field analysis for the asymptotic behavior of deep BitNet-like architectures as smooth quantization parameters approach zero. We establish that empirical measures of latent weights converge weakly to solutions of…

Optimization and Control · Mathematics 2025-09-03 Dongwon Kim , Dongseok Lee

We treat the heat equation with singular drift terms and its generalization: the linearized Navier-Stokes system. In the first case, we obtain boundedness of weak solutions for highly singular, "supercritical" data. In the second case, we…

Analysis of PDEs · Mathematics 2019-06-03 Qi S Zhang