Related papers: Pointwise gradient estimate of the ritz projection
In this paper, we consider a class of finite-sum convex optimization problems whose objective function is given by the summation of $m$ ($\ge 1$) smooth components together with some other relatively simple terms. We first introduce a…
We investigate the enumerative geometry of point configurations in projective space. We define "projective configuration counts": these enumerate configurations of points in projective space such that certain specified subsets are in fixed…
A subgradient method is presented for solving general convex optimization problems, the main requirement being that a strictly-feasible point is known. A feasible sequence of iterates is generated, which converges to within user-specified…
For the reference triangle or tetrahedron $T$, we study the stability properties of the $L^2(T)$-projection $\Pi_N$ onto the space of polynomials of degree $N$. We show $\|\Pi_N u\|_{L^2(\partial T)}^2 \leq C \|u\|_{L^2(T)} \|u\|_{H^1(T)}$.…
We prove the weak and the strong convergence of the trajectories of the continuous gradient projection method under some mild assumptions on the objective function and the step size function. Moreover, we estimate the decay rate to…
This paper deals with composite optimization problems having the objective function formed as the sum of two terms, one has Lipschitz continuous gradient along random subspaces and may be nonconvex and the second term is simple and…
The {\it number rigidity} of a stationary point process $\mathsf{P}$ entails that for a bounded set $A$ the knowledge of $\mathsf{P}$ on $A^{c}$ a.s. determines $\mathsf{P}(A)$; the $k$-order rigidity means the moments of $\mathsf{P}1_{A}$…
We provide novel theoretical results regarding local optima of regularized $M$-estimators, allowing for nonconvexity in both loss and penalty functions. Under restricted strong convexity on the loss and suitable regularity conditions on the…
$D$-optimal designs originate in statistics literature as an approach for optimal experimental designs. In numerical analysis points and weights resulting from maximal determinants turned out to be useful for quadrature and interpolation.…
We introduce a polynomial time algorithm for optimizing the class of star-convex functions, under no restrictions except boundedness on a region about the origin, and Lebesgue measurability. The algorithm's performance is polynomial in the…
This paper focuses on the problem of \emph{constrained} \emph{stochastic} optimization. A zeroth order Frank-Wolfe algorithm is proposed, which in addition to the projection-free nature of the vanilla Frank-Wolfe algorithm makes it gradient…
We propose a simple and effective method for designing approximation formulas for weighted analytic functions. We consider spaces of such functions according to weight functions expressing the decay properties of the functions. Then, we…
We study multipoint Pad\'e approximants of type $(n,n)$ for the Hurwitz zeta function $f(a)=\zeta(s,a)$ with $\Re s>1$, constructed at quantile nodes $a_{n,j}=n\alpha_{n,j}$ generated by a real-analytic density $\kappa$ on…
Stochastic convex optimization is a basic and well studied primitive in machine learning. It is well known that convex and Lipschitz functions can be minimized efficiently using Stochastic Gradient Descent (SGD). The Normalized Gradient…
Bounds on the log partition function are important in a variety of contexts, including approximate inference, model fitting, decision theory, and large deviations analysis. We introduce a new class of upper bounds on the log partition…
The problem of the minimization of least squares functionals with $\ell^1$ penalties is considered in an infinite dimensional Hilbert space setting. While there are several algorithms available in the finite dimensional setting there are…
We consider the Hardy-Littlewood maximal function associated with ball averages on spaces with exponential volume growth. We focus on discrete groups with balls defined by invariant metrics associated with a variety of length functions.…
We investigate the properties of the simultaneous projection method as applied to countably infinitely many closed and linear subspaces of a real Hilbert space. We establish the optimal error bound for linear convergence of this method,…
Minimizing a convex function of a measure with a sparsity-inducing penalty is a typical problem arising, e.g., in sparse spikes deconvolution or two-layer neural networks training. We show that this problem can be solved by discretizing the…
We introduce a new surrogate loss function called orbit loss in the structured prediction framework, which has good theoretical and practical advantages. While the orbit loss is not convex, it has a simple analytical gradient and a simple…