Related papers: Further Improvements to the Lower Bound for an Aut…
Deriving sharp and computable upper bounds of the Lipschitz constant of deep neural networks is crucial to formally guarantee the robustness of neural-network based models. We analyse three existing upper bounds written for the $l^2$ norm.…
We design an algorithm which finds an $\epsilon$-approximate stationary point (with $\|\nabla F(x)\|\le \epsilon$) using $O(\epsilon^{-3})$ stochastic gradient and Hessian-vector products, matching guarantees that were previously available…
One of the hard optimization problems that has a semi-definite relaxation with quantitative bound on the approximation error is the maximization of a convex quadratic form on the hypercube. The relaxation not only yields an upper bound on…
The well-known Bennett-Hoeffding bound for sums of independent random variables is refined, by taking into account truncated third moments, and at that also improved by using, instead of the class of all increasing exponential functions,…
We study realizable continual linear regression under random task orderings, a common setting for developing continual learning theory. In this setup, the worst-case expected loss after $k$ learning iterations admits a lower bound of…
We present an adaptive trust-region method for unconstrained optimization that allows inexact solutions to the trust-region subproblems. Our method is a simple variant of the classical trust-region method of \citet{sorensen1982newton}. The…
We consider the averages of a function $ f$ on $ \mathbb R ^{n}$ over spheres of radius $ 0< r< \infty $ given by $ A_{r} f (x) = \int_{\mathbb S ^{n-1}} f (x-r y) \; d \sigma (y)$, where $ \sigma $ is the normalized rotation invariant…
This paper studies the complexity of finding an $\epsilon$-stationary point for stochastic bilevel optimization when the upper-level problem is nonconvex and the lower-level problem is strongly convex. Recent work proposed the first-order…
Error bounds are central objects in optimization theory and its applications. They were for a long time restricted only to the theory before becoming over the course of time a field of itself. This paper is devoted to the study of error…
This paper provides inference methods for best linear approximations to functions which are known to lie within a band. It extends the partial identification literature by allowing the upper and lower functions defining the band to be any…
In this paper, we study the proximal incremental aggregated gradient(PIAG) algorithm for minimizing the sum of L-smooth nonconvex component functions and a proper closed convex function. By exploiting the L-smooth property and with the help…
Let P be a set of points and $L$ a set of lines in (F_p)^2, with |P|,|L|\leq N and N<p. We show that P and L generate no more than C N^(3/2 - 1/806 + o(1)) incidences for some absolute constant C. This improves by an order of magnitude on…
We present two approximate versions of the proximal subgradient method for minimizing the sum of two convex functions (not necessarily differentiable). The algorithms involve, at each iteration, inexact evaluations of the proximal operator…
For all functions on an arbitrary open set $\Omega\subset\R^3$ with zero boundary values, we prove the optimal bound \[ \sup_{\Omega}|u| \leq (2\pi)^{-1/2} \left(\int_{\Omega}|\nabla u|^2 \,dx\, \int_{\Omega}|\Delta u|^2 \,dx\right)^{1/4}.…
Conditional Gradient algorithms (aka Frank-Wolfe algorithms) form a classical set of methods for constrained smooth convex minimization due to their simplicity, the absence of projection steps, and competitive numerical performance. While…
We consider an overdetermined problem for Laplace equation on a disk with partial boundary data where additional pointwise data inside the disk have to be taken into account. After reformulation, this ill-posed problem reduces to a bounded…
The (Non-Preemptive) Throughput Maximization problem is a natural and fundamental scheduling problem. We are given $n$ jobs, where each job $j$ is characterized by a processing time and a time window, contained in a global interval $[0,T)$,…
In this work, we deal with approximations for distribution functions of non-negative random variables. More specifically, we construct continuous approximants using an acceleration technique over a well-know inversion formula for Laplace…
We consider structured optimisation problems defined in terms of the sum of a smooth and convex function, and a proper, l.s.c., convex (typically non-smooth) one in reflexive variable exponent Lebesgue spaces $L_{p(\cdot)}(\Omega)$. Due to…
We develop minimax optimal risk bounds for the general learning task consisting in predicting as well as the best function in a reference set G up to the smallest possible additive term, called the convergence rate. When the reference set…