Related papers: Constructing a subgradient from directional deriva…
Fitting a function by using linear combinations of a large number $N$ of `simple' components is one of the most fruitful ideas in statistical learning. This idea lies at the core of a variety of methods, from two-layer neural networks to…
Motivated by extending the functional stochastic calculus, to important functionals to which it does not apply, a notion of functional derivative along a curve is introduced. This new setting is developed by incorporating path-dependent…
This paper presents a stochastic block-coordinate proximal Newton method for minimizing the sum of a blockwise Lipschitz-continuously differentiable function and a separable nonsmooth convex function. At each iteration, the method randomly…
In this paper we present a variant of the proximal forward-backward splitting iteration for solving nonsmooth optimization problems in Hilbert spaces, when the objective function is the sum of two nondifferentiable convex functions. The…
In this paper, we address stochastic optimization problems involving a composition of a non-smooth outer function and a smooth inner function, a formulation frequently encountered in machine learning and operations research. To deal with…
In this work, we construct a proximal average for two prox-bounded functions, which recovers the classical proximal average for two convex functions. The new proximal average transforms continuously in epi-topology from one proximal hull to…
It is shown that if $\gamma$ is a path of finite $p$ variation ($1\leq p< 2$) in a euclidean vector space and $f,g,h$ are Lipschitz functions on the trace of $\gamma$ then $s\mapsto F(s)=\int_\gamma f^sg dh$ defines an entire holomorphic…
Let $p$ be a polynomial in the non-commuting variables $(a,x)=(a_1,...,a_{g_a},x_1,...,x_{g_x})$. If $p$ is convex in the variables $x$, then $p$ has degree two in $x$ and moreover, $p$ has the form $p = L + \Lambda ^T \Lambda,$ where $L$…
Using backpropagation to compute gradients of objective functions for optimization has remained a mainstay of machine learning. Backpropagation, or reverse-mode differentiation, is a special case within the general family of automatic…
In this paper we study the radial epiderivative notion for nonconvex functions, which extends the (classical) directional derivative concept. The paper presents new definition and new properties for this notion and establishes relationships…
In this paper we develop a geometric approach to convex subdifferential calculus in finite dimensions with employing some ideas of modern variational analysis. This approach allows us to obtain natural and rather easy proofs of basic…
We show that a differentiable function on the 2-Wasserstein space is geodesically convex if and only if it is also convex along a larger class of curves which we call `acceleration-free'. In particular, the set of acceleration-free curves…
With the increasing interest in applying the methodology of difference-of-convex (dc) optimization to diverse problems in engineering and statistics, this paper establishes the dc property of many well-known functions not previously known…
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a…
The recent results of An, Luan, and Yen [Differential stability in convex optimization via generalized polyhedrality. Vietnam J. Math. https://-doi.org/10.1007/s10013-024-00721-y] on differential stability of parametric optimization…
It is hereby established that the set of Lipschitz functions $f:\mathcal{U}\rightarrow \mathbb{R}$ ($\mathcal{U}$ nonempty open subset of $\ell_{d}^{1}$) with maximal Clarke subdifferential contains a linear subspace of uncountable…
We study the oracle complexity of producing $(\delta,\epsilon)$-stationary points of Lipschitz functions, in the sense proposed by Zhang et al. [2020]. While there exist dimension-free randomized algorithms for producing such points within…
The usual approach to developing and analyzing first-order methods for smooth convex optimization assumes that the gradient of the objective function is uniformly smooth with some Lipschitz constant $L$. However, in many settings the…
In this work, several sharp bounds for the \v{C}eby\v{s}ev functional involving various type of functions are proved. In particular, for the \v{C}eby\v{s}ev functional of two absolutely continuous functions whose first derivatives are both…
The paper addresses the study and applications of a broad class of extended-real-valued functions, known as optimal value or marginal functions, which are frequently appeared in variational analysis, parametric optimization, and a variety…