Related papers: What is the gradient of a scalar function of a sym…
Motivated by the refinements and reverses of arithmetic-geometric mean and arithmetic-harmonic mean inequalities for scalars and matrices, in this article, we generalize the scalar and matrix inequalities for the difference between…
Fractional derivatives are a well-studied generalization of integer order derivatives. Naturally, for optimization, it is of interest to understand the convergence properties of gradient descent using fractional derivatives. Convergence…
Simplex gradients are an essential feature of many derivative free optimization algorithms, and can be employed, for example, as part of the process of defining a direction of search, or as part of a termination criterion. The calculation…
Subdifferentials (in the sense of convex analysis) of matrix-valued functions defined on $\mathbb{R}^d$ that are convex with respect to the L\"{o}wner partial order can have a complicated structure and might be very difficult to compute…
Spectral functions of symmetric matrices -- those depending on matrices only through their eigenvalues -- appear often in optimization. A cornerstone variational analytic tool for studying such functions is a formula relating their…
In this paper we focus on the linear functionals defining an approximate version of the gradient of a function. These functionals are often used when dealing with optimization problems where the computation of the gradient of the objective…
The simple product formulae for derivatives of scalar functions raised to different powers are generalized for functions which take values in the set of symmetric positive definite matrices. These formulae are fundamental in derivation of…
We define a symmetric derivative on an arbitrary nonempty closed subset of the real numbers and derive some of its properties. It is shown that real-valued functions defined on time scales that are neither delta nor nabla differentiable can…
In this preliminary study, we provide two methods for estimating the gradients of functions of real value. Both methods are built on derivative estimations that are calculated using the standard method or the Squire-Trapp method for any…
We investigate the computation of the gradient of the value function in parametric convex optimization problems. We derive general expression for the gradient of the value function in terms of the cost function, constraints and Lagrange…
Factorization-based gradient descent is a scalable and efficient algorithm for solving low-rank matrix completion. Recent progress in structured non-convex optimization has offered global convergence guarantees for gradient descent under…
For Dirac operators, which have discrete spectra, the concept of eigenvalues gradient is given and formulae for this gradients are obtained in terms of normalized eigenfunctions. It is shown how the gradient is being used to describe…
A symmetric function of $N$ variables can be given in terms of symmetric polynomials of these variables. We determine those symmetric polynomials in which the dual differential operators take the neatest form when expressed in terms of our…
This paper aims at achieving a "good" estimator for the gradient of a function on a high-dimensional space. Often such functions are not sensitive in all coordinates and the gradient of the function is almost sparse. We propose a method for…
A generalized matrix function $d_\chi^G : M_n(\mathbb{C}) \rightarrow \mathbb{C}$ is a function constructed by a subgroup $G$ of $S_n$ and a complex valued function $\chi$ of $G$. The main purpose of this paper is to find a necessary and…
We study gradient testing and gradient estimation of smooth functions using only a comparison oracle that, given two points, indicates which one has the larger function value. For any smooth $f\colon\mathbb R^n\to\mathbb R$,…
Gradient approximations are a class of numerical approximation techniques that are of central importance in numerical optimization. In derivative-free optimization, most of the gradient approximations, including the simplex gradient,…
Low-rank matrix estimation plays a central role in various applications across science and engineering. Recently, nonconvex formulations based on matrix factorization are provably solved by simple gradient descent algorithms with strong…
For a locally Lipschitz continuous function $f:X\to\mathbb{R}$ the generalized gradient $\partial f(x)$ of Clarke is used to develop some (set-valued) gradient on a set $A\subset X$. Existence, uniqueness and some approximation are…
The paper looks at a scaled variant of the stochastic gradient descent algorithm for the matrix completion problem. Specifically, we propose a novel matrix-scaling of the partial derivatives that acts as an efficient preconditioning for the…