Related papers: Convergence to minima for the continuous version o…
This work considers the question: what convergence guarantees does the stochastic subgradient method have in the absence of smoothness and convexity? We prove that the stochastic subgradient method, on any semialgebraic locally Lipschitz…
We focus on a sequence of functions $\{f_n\}$, defined on a compact manifold with boundary $S$, converging in the $C^k$ metric to a limit $f$. A common assumption implicitly made in the empirical sciences is that when such functions…
Here is one of the results of this paper (with the convention ${{1}\over {0}}=+\infty$): Let $X$ be a real Hilbert space and let $J:X\to {\bf R}$ be a $C^1$ functional, with compact derivative, such that $$\alpha^*:=\max\left…
We prove that, for any closed semialgebraic subset $W$ of $\mathbb{R}^n$ and for any positive integer $p$, there exists a Nash function $f:\mathbb{R}^n\setminus W\longrightarrow (0, \infty)$ which is equivalent to the distance function from…
In this paper, we point out a very flexible scheme within which a strict minimax inequality occurs. We then show the fruitfulness of this approach presenting a series of various consequences. Here is one of them: Let $Y$ be a…
We prove the exact worst-case convergence rate of gradient descent for smooth strongly convex optimization, with respect to the performance criterion $\Vert \nabla f(x_N)\Vert^2/(f(x_0)-f_*)$. The proof differs from the previous one by…
One means of fitting functions to high-dimensional data is by providing smoothness constraints. Recently, the following smooth function approximation problem was proposed: given a finite set $E \subset \mathbb{R}^d$ and a function $f: E…
This paper develops and analyzes an accelerated proximal descent method for finding stationary points of nonconvex composite optimization problems. The objective function is of the form $f+h$ where $h$ is a proper closed convex function,…
We consider the minimization of non-convex quadratic forms regularized by a cubic term, which exhibit multiple saddle points and poor local minima. Nonetheless, we prove that, under mild assumptions, gradient descent approximates the…
Let $\Omega$ be an open subset of $\mathbb R^n$, and let $f: \Omega \to \mathbb R$ be differentiable $\mathcal H^k$-almost everywhere, for some nonnegative integer $k < n$, where $\mathcal H^k$ denotes the $k$=dimensional Hausdorff measure.…
In this paper, we establish an improved version of a saddle point theorem ([4]) removing a weak lower semicontinuity assumption at all. We then revisit some of the applications of that theorem in the light of such an improvement. For…
We show that Lipschitz solutions $u$ of $\mathrm{div}\, G(\nabla u)=0$ in $B_1\subset\mathbb R^2$ are $C^1$, for strictly monotone vector fields $G\in C^0(\mathbb R^2;\mathbb R^2)$ satisfying a mild ellipticity condition. If $G=\nabla F$…
We prove the exact worst-case convergence rate of gradient descent for smooth strongly convex optimization on $\mathbb{R}^d$. Concretely, assuming that the objective function $f$ is $\mu$-strongly convex and $L$-smooth, we identify the…
Fix a constant $0<\alpha <1$. For a $C^1$ function $f:\mathbb{R}^k\rightarrow \mathbb{R}$, a point $x$ and a positive number $\delta >0$, we say that Armijo's condition is satisfied if $f(x-\delta \nabla f(x))-f(x)\leq -\alpha \delta…
Composite optimization problems, where the sum of a smooth and a merely lower semicontinuous function has to be minimized, are often tackled numerically by means of proximal gradient methods as soon as the lower semicontinuous part of the…
Consider the consensus problem of minimizing $f(x)=\sum_{i=1}^n f_i(x)$ where each $f_i$ is only known to one individual agent $i$ out of a connected network of $n$ agents. All the agents shall collaboratively solve this problem and obtain…
Classical results show that gradient descent converges linearly to minimizers of smooth strongly convex functions. A natural question is whether there exists a locally nearly linearly convergent method for nonsmooth functions with quadratic…
We present here a new method for approximating functions defined on superreflexive Banach spaces by differentiable functions with $\alpha$-H\"older derivatives (for some $0<\alpha\leq 1$). The smooth approximation is given by means of an…
Let $E \subset \mathbb{R}^n$ be a compact set, and $f:E \to \mathbb{R}$. How can we tell if there exists a convex extension $F \in C^{1,1}(\mathbb{R}^n)$ of $f$, i.e. satisfying $F|_E = f|_E$? Assuming such an extension exists, how small…
The usual approach to developing and analyzing first-order methods for non-smooth (stochastic or deterministic) convex optimization assumes that the objective function is uniformly Lipschitz continuous with parameter $M_f$. However, in many…