English

Zeroth-order Gradient and Quasi-Newton Methods for Nonsmooth Nonconvex Stochastic Optimization

Optimization and Control 2025-10-21 v3

Abstract

We consider the minimization of a Lipschitz continuous and expectation-valued function, denoted by ff and defined as f(x)E[f~(x,ξ)]f(\mathbf{x}) \triangleq \mathbb{E}[\tilde{f}(\mathbf{x}, \mathbf{\xi})], over a closed and convex set X\mathcal{X}. We obtain asymptotics as well as rate and complexity guarantees for computing approximate Clarke-stationary points via zeroth-order schemes. We adopt an approach reliant on minimizing fηf_{\eta} where fη(x)Eu[x,f(x+ηu)]f_{\eta}(\mathbf{x}) \triangleq \mathbb{E}_{\mathbf{u}}\left[\mathbf{x}, f(\mathbf{x}+\eta \mathbf{u})\, \right], u\mathbf{u} is a random variable defined on a unit sphere, and η>0\eta > 0. In fact, it is known that a stationary point of the η\eta-smoothed problem is an η\eta-stationary point for the original problem in the Clarke sense. In such a setting, we develop two schemes with promising empirical behavior. (I) We develop a variance-reduced zeroth-order gradient framework (VRG-ZO) for minimizing fηf_{\eta} over X\mathcal{X}. In this setting, we make two sets of contributions for the sequence generated by the proposed zeroth-order gradient scheme. (a) The residual function of the smoothed problem tends to zero almost surely along the generated sequence, guaranteeing η\eta-Clarke stationary solutions of the original problem; (b) To compute an x\mathbf{x} such that the expected norm of the residual of the η\eta-smoothed problem is within ϵ\epsilon requires no greater than O(n1/2(L0η1+L02)ϵ2)\mathcal{O}({n^{1/2}}{(L_0\eta^{-1} +L_0^2)} \epsilon^{-2}) projection steps and O(n3/2(L03η2+L05)ϵ4)\mathcal{O}({n^{3/2}(L_0^3\eta^{-2}+L_0^5)} \epsilon^{-4}) function evaluations. (II) Our second scheme is a zeroth-order stochastic quasi-Newton scheme (VRSQN-ZO) reliant on randomized and Moreau smoothing; the iteration and sample complexities are O(L04n2η4ϵ2)\mathcal{O}({L_0^{4}}{n^{2}}{\eta^{-4}}\epsilon^{-2}) and O(L09n5η5ϵ5)\mathcal{O}(L_0^{9} n^{5}\eta^{-5}\epsilon^{-5}), respectively.

Keywords

Cite

@article{arxiv.2401.08665,
  title  = {Zeroth-order Gradient and Quasi-Newton Methods for Nonsmooth Nonconvex Stochastic Optimization},
  author = {Luke Marrinan and Uday V. Shanbhag and Farzad Yousefian},
  journal= {arXiv preprint arXiv:2401.08665},
  year   = {2025}
}
R2 v1 2026-06-28T14:18:29.466Z