Related papers: Convergence to minima for the continuous version o…
In this paper, using the theory developed in [8], we obtain some results of a totally new type about a class of non-local problems. Here is a sample: Let $\Omega\subset {\bf R}^n$ be a smooth bounded domain, with $n\geq 4$, let $a, b,…
In nonsmooth optimization, a negative subgradient is not necessarily a descent direction, making the design of convergent descent methods based on zeroth-order and first-order information a challenging task. The well-studied bundle methods…
A hypodifferential is a compact family of affine mappings that defines a local max-type approximation of a nonsmooth convex function. We present a general theory of hypodifferentials of nonsmooth convex functions defined on a Banach space.…
In this paper, we approach the problem of finding the zeros of the sum of a maximally monotone operator and a monotone and Lipschitz continuous one in a real Hilbert space via an implicit forward-backward-forward dynamical system with…
While much progress has been achieved over the last decades in neuro-inspired machine learning, there are still fundamental theoretical problems in gradient-based learning using combinations of neurons. These problems, such as saddle points…
We consider first-order methods with constant step size for minimizing locally Lipschitz coercive functions that are tame in an o-minimal structure on the real field. We prove that if the method is approximated by subgradient trajectories,…
This note corrects a gap and improves results in an earlier paper by the first named author. More precisely, it is shown that on weakly compactly generated Banach spaces X which admit a C^{p} smooth norm, one can uniformly approximate…
Many machine learning and data science tasks require solving non-convex optimization problems. When the loss function is a sum of multiple terms, a popular method is the stochastic gradient descent. Viewed as a process for sampling the loss…
We study the convergence properties of gradient descent for training deep linear neural networks, i.e., deep matrix factorizations, by extending a previous analysis for the related gradient flow. We show that under suitable conditions on…
In this paper we develop some new techniques to study the multiscale elliptic equations in the form of $-\text{div} \big(A_\varepsilon \nabla u_{\varepsilon} \big) = 0$, where $A_\varepsilon(x) = A(x, x/\varepsilon_1,\cdots,…
The gradient method for minimize a differentiable convex function on Riemannian manifolds with lower bounded sectional curvature is analyzed in this paper. The analysis of the method is presented with three different finite procedures for…
Based on a result by Taylor, Hendrickx, and Glineur (J. Optim. Theory Appl., 178(2):455--476, 2018) on the attainable convergence rate of gradient descent for smooth and strongly convex functions in terms of function values, an elementary…
In this paper, we present a new complexity result for the gradient descent method with an appropriately fixed stepsize for minimizing a strongly convex function with locally $\alpha$-H{\"o}lder continuous gradients ($0 < \alpha \leq 1$).…
Let $f$ be a nonnegative function of class $C^k$ ($k \geq 2$) such that $f^{(k)}$ is H\''older continuous with exponent $\alpha$ in $(0,1]$. If $f'(x) = \cdots = f^{(k)}(x) = 0$ when $f(x) = 0$, we show that $f^{\mu}$ is differentiable for…
In a Hilbert space setting $\mathcal H$, we study the fast convergence properties as $t \to + \infty$ of the trajectories of the second-order differential equation $ \ddot{x}(t) + \frac{\alpha}{t} \dot{x}(t) + \nabla \Phi (x(t)) = g(t)$,…
This work provides the first convergence analysis for the Randomized Block Coordinate Descent method for minimizing a function that is both H\"older smooth and block H\"older smooth. Our analysis applies to objective functions that are…
Let $U\subseteq\mathbb{R}^{n}$ be open and convex. We show that every (not necessarily Lipschitz or strongly) convex function $f:U\to\mathbb{R}$ can be approximated by real analytic convex functions, uniformly on all of $U$. In doing so we…
We extend deconvolution in a periodic setting to deal with functional data. The resulting functional deconvolution model can be viewed as a generalization of a multitude of inverse problems in mathematical physics where one needs to recover…
The goal of the paper is to study the particular class of regularly ${\mathcal{H}}$-convex functions, when ${\mathcal{H}}$ is the set ${\mathcal{L}\widehat{C}}(X,{\mathbb{R}})$ of real-valued Lipschitz continuous classically concave…
We remove the dependence on the `hot-spots' conjecture in two of the main theorems of the recent paper of Nickl (2024, Annals of Statistics). Specifically, we characterise the minimax convergence rates for estimation of the transition…