Related papers: On the diameter of subgradient sequences in o-mini…
In this paper, we suggest a new framework for analyzing primal subgradient methods for nonsmooth convex optimization problems. We show that the classical step-size rules, based on normalization of subgradient, or on the knowledge of optimal…
In stochastic convex optimization problems, most existing adaptive methods rely on prior knowledge about the diameter bound $D$ when the smoothness or the Lipschitz constant is unknown. This often significantly affects performance as only a…
This paper has two themes that are intertwined: The first is the dynamics of certain piecewise affine maps on the Euclidean space that arise from a class of analog-to-digital conversion methods called Sigma-Delta quantization. The second is…
Stochastic optimization lies at the core of most statistical learning models. The recent great development of stochastic algorithmic tools focused significantly onto proximal gradient iterations, in order to find an efficient approach for…
In previous papers by A. Kameyama and by J. Kigami distances on fractals have been discussed having two different but similar properties. One property is that the maps defining the fractal are Lipschitz of prescribed constants less than 1,…
We consider the problem of minimizing a convex objective which is the sum of a smooth part, with Lipschitz continuous gradient, and a nonsmooth part. Inspired by various applications, we focus on the case when the nonsmooth part is a…
Finite subset spaces of a metric space $X$ form a nested sequence under natural isometric embeddings $X=X(1)\subset X(2)\subset\dots$. We prove that this sequence admits Lipschitz retractions $X(n)\to X(n-1)$ when $X$ is a Hilbert space.
In this paper, we study the shift on the space of uniformly bounded continuous functions band-limited in a given compact interval with the standard topology of tempered distributions. We give a constructive proof of the existence of minimal…
We propose a mini-batching scheme for improving the theoretical complexity and practical performance of semi-stochastic gradient descent applied to the problem of minimizing a strongly convex composite function represented as the sum of an…
Given an open subset $\Omega$ of a Banach space and a Lipschitz function $u_0: \overline{\Omega} \to \mathbb{R},$ we study whether it is possible to approximate $u_0$ uniformly on $\Omega$ by $C^k$-smooth Lipschitz functions which coincide…
This paper gives a unified and succinct approach to the $O(1/\sqrt{k}), O(1/k),$ and $O(1/k^2)$ convergence rates of the subgradient, gradient, and accelerated gradient methods for unconstrained convex minimization. In the three cases the…
We show that the log-likelihood of several probabilistic graphical models is Lipschitz continuous with respect to the lp-norm of the parameters. We discuss several implications of Lipschitz parametrization. We present an upper bound of the…
Let $\Lambda$ be a uniformly discrete set and $S$ be a compact set in $R$. We prove that if there exists a bounded sequence of functions in Paley--Wiener space $PW_S$, which approximates $\delta-$functions on $\Lambda$ with $l^2-$error $d$,…
Adaptive gradient methods are typically used for training over-parameterized models. To better understand their behaviour, we study a simplistic setting -- smooth, convex losses with models over-parameterized enough to interpolate the data.…
The performance of Metropolis-Hastings algorithms is highly sensitive to the choice of step size, and miss-specification can lead to severe loss of efficiency. We study algorithms with randomized step sizes, considering both…
In this paper we propose a distributed version of a randomized block-coordinate descent method for minimizing the sum of a partially separable smooth convex function and a fully separable non-smooth convex function. Under the assumption of…
This work provides the first convergence analysis for the Randomized Block Coordinate Descent method for minimizing a function that is both H\"older smooth and block H\"older smooth. Our analysis applies to objective functions that are…
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a…
This paper discusses several (sub)gradient methods attaining the optimal complexity for smooth problems with Lipschitz continuous gradients, nonsmooth problems with bounded variation of subgradients, weakly smooth problems with H\"older…
In this paper, we study Multi-$\mathcal{K}$-equivalence of multi-germs of functions on the plane, definable in a polynomially bounded o-minimal structure. We partition the germ of the plane at origin into zones of arcs in such a way that it…