English
Related papers

Related papers: Gradient descent in a generalised Bregman distance…

200 papers

Lipschitz continuity of the gradient mapping of a continuously differentiable function plays a crucial role in designing various optimization algorithms. However, many functions arising in practical applications such as low rank matrix…

Optimization and Control · Mathematics 2020-12-25 Mahesh Chandra Mukkamala , Jalal Fadili , Peter Ochs

The paper introduces scaled Bregman distances of probability distributions which admit non-uniform contributions of observed events. They are introduced in a general form covering not only the distances of discrete and continuous stochastic…

Information Theory · Computer Science 2021-05-12 Wolfgang Stummer , Igor Vajda

We propose two numerical algorithms in the fully nonconvex setting for the minimization of the sum of a smooth function and the composition of a nonsmooth function with a linear operator. The iterative schemes are formulated in the spirit…

Optimization and Control · Mathematics 2020-08-03 Radu Ioan Bot , Dang-Khoa Nguyen

We study two variants of the mirror descent-ascent (MDA) algorithm for solving min-max problems on the space of measures: simultaneous and alternating. We work under assumptions of convexity-concavity and relative smoothness of the payoff…

Optimization and Control · Mathematics 2026-05-08 Razvan-Andrei Lascu , Mateusz B. Majka , Łukasz Szpruch

In this paper, we generalize the notions of centroids and barycenters to the broad class of information-theoretic distortion measures called Bregman divergences. Bregman divergences are versatile, and unify quadratic geometric distances…

Computational Geometry · Computer Science 2007-11-22 Frank Nielsen , Richard Nock

In this paper, we consider a class of generalized difference-of-convex functions (DC) programming, whose objective is the difference of two convex (not necessarily smooth) functions plus a decomposable (possibly nonconvex) function with…

Optimization and Control · Mathematics 2024-09-10 Chenjian Pan , Yingxin Zhou , Hongjin He , Chen Ling

Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have been widely used in modern machine learning. Here, we study the…

Machine Learning · Computer Science 2022-08-31 Xiaoxia Wu , Edgar Dobriban , Tongzheng Ren , Shanshan Wu , Zhiyuan Li , Suriya Gunasekar , Rachel Ward , Qiang Liu

Nonconvex-nonconcave saddle-point optimization in machine learning has triggered lots of research for studying non-monotone variational inequalities (VI). In this work, we introduce two mirror frameworks, called mirror extragradient method…

Optimization and Control · Mathematics 2023-01-02 Hui Zhang , Yu-Hong Dai

The (global) Lipschitz smoothness condition is crucial in establishing the convergence theory for most optimization methods. Unfortunately, most machine learning and signal processing problems are not Lipschitz smooth. This motivates us to…

Optimization and Control · Mathematics 2019-04-23 Qiuwei Li , Zhihui Zhu , Gongguo Tang , Michael B. Wakin

We study the convergence rate of Bregman gradient methods for convex optimization in the space of measures on a $d$-dimensional manifold. Under basic regularity assumptions, we show that the suboptimality gap at iteration $k$ is in…

Optimization and Control · Mathematics 2023-03-15 Lénaïc Chizat

Modern statistical applications often involve minimizing an objective function that may be nonsmooth and/or nonconvex. This paper focuses on a broad Bregman-surrogate algorithm framework including the local linear approximation, mirror…

Optimization and Control · Mathematics 2021-12-20 Yiyuan She , Zhifeng Wang , Jiuwu Jin

We present a family of algorithms, called descent algorithms, for optimizing convex and non-convex functions. We also introduce a new first-order algorithm, called rescaled gradient descent (RGD), and show that RGD achieves a faster…

Optimization and Control · Mathematics 2020-01-07 Ashia Wilson , Lester Mackey , Andre Wibisono

Many metric learning tasks, such as triplet learning, nearest neighbor retrieval, and visualization, are treated primarily as embedding tasks where the ultimate metric is some variant of the Euclidean distance (e.g., cosine or Mahalanobis),…

Machine Learning · Computer Science 2023-11-22 Fred Lu , Edward Raff , Francis Ferraro

We consider a variable metric linesearch based proximal gradient method for the minimization of the sum of a smooth, possibly nonconvex function plus a convex, possibly nonsmooth term. We prove convergence of this iterative algorithm to a…

Numerical Analysis · Mathematics 2017-04-11 Silvia Bonettini , Ignace Loris , Federica Porta , Marco Prato , Simone Rebegoldi

The mirror descent algorithm (MDA) generalizes gradient descent by using a Bregman divergence to replace squared Euclidean distance. In this paper, we similarly generalize the alternating direction method of multipliers (ADMM) to Bregman…

Optimization and Control · Mathematics 2014-07-09 Huahua Wang , Arindam Banerjee

The aim of this paper is to provide an overview of recent development related to Bregman distances outside its native areas of optimization and statistics. We discuss approaches in inverse problems and image processing based on Bregman…

Optimization and Control · Mathematics 2015-05-21 Martin Burger

Stochastic optimization powers the scalability of modern artificial intelligence, spanning machine learning, deep learning, reinforcement learning, and large language model training. Yet, existing theory remains largely confined to Hilbert…

Machine Learning · Computer Science 2025-09-18 Johnny R. Zhang , Xiaomei Mi , Gaoyuan Du , Qianyi Sun , Shiqi Wang , Jiaxuan Li , Wenhua Zhou

Recent work by Woodworth et al. (2020) shows that the optimization dynamics of gradient descent for overparameterized problems can be viewed as low-dimensional dual dynamics induced by a mirror map, explaining the implicit regularization…

Machine Learning · Computer Science 2024-10-21 Shuyang Wang , Diego Klabjan

Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking. These phenomena appear across architectures -- in…

Machine Learning · Computer Science 2026-01-01 Alan Oursland

This paper proposes and develops new linesearch methods with inexact gradient information for finding stationary points of nonconvex continuously differentiable functions on finite-dimensional spaces. Some abstract convergence results for a…

Optimization and Control · Mathematics 2023-01-03 Pham Duy Khanh , Boris S. Mordukhovich , Dat Ba Tran