Related papers: A posteriori superlinear convergence bounds for bl…
We introduce a framework to accelerate the convergence of gradient-based methods with online learning. The framework learns to scale the gradient at each iteration through an online learning algorithm and provably accelerates gradient-based…
The aim of this paper is to give a short overview on error bounds and to provide the first bricks of a unified theory. Inspired by the works of [8, 15, 13, 16, 10], we show indeed the centrality of the Lojasiewicz gradient inequality. For…
In this paper we prove complex bounds, also referred to as a priori bounds, for real analytic (and even C3) interval maps. This means that we associate to such a map a complex box mapping (which provides a kind of Markov structure),…
The variation of spectral subspaces for linear self-adjoint operators under an additive bounded off-diagonal perturbation is studied. To this end, the optimization approach for general perturbations in [J. Anal. Math., to appear;…
Gradient descent and stochastic gradient descent are central to modern machine learning, yet their behavior under large step sizes remains theoretically unclear. Recent work suggests that acceleration often arises near the edge of…
Deep learning has aroused extensive attention due to its great empirical success. The efficiency of the block coordinate descent (BCD) methods has been recently demonstrated in deep neural network (DNN) training. However, theoretical…
This article derives lower bounds on the convergence rate of continuous-time gradient-based optimization algorithms. The algorithms are subjected to a time-normalization constraint that avoids a reparametrization of time in order to make…
The nonlinear conjugate gradient methods are known to be an effective approach for standard unconstrained optimization problems especially for large-scale problems. This paper proposes a proximal nonlinear conjugate gradient method, which…
A new error bound for the linear complementarity problem is given when the involved matrix is a B-matrix. It is shown that this bound is sharper than some previous bounds [C.Q. Li, Y.T. Li. Note on error bounds for linear complementarity…
This paper directly builds upon previous work where we introduced new reduced basis a posteriori error bounds for parametrized saddle point problems based on Brezzi's theory. We here sharpen these estimates for the special case of a…
We prove an upper bound on the degree complexity of Putinar's Positivstellensatz. This bound is much worse than the one obtained previously for Schm\"udgen's Positivstellensatz but it depends on the same parameters. As a consequence, we get…
This paper is concerned with the recovery of (approximate) solutions to parabolic problems from incomplete and possibly inconsistent observational data, given on a time-space cylinder that is a strict subset of the computational domain…
We consider finite element solutions to optimization problems, where the state depends on the possibly constrained control through a linear partial differential equation. Basing upon a reduced and rescaled optimality system, we derive a…
Aimed at explaining the surprisingly good generalization behavior of overparameterized deep networks, recent works have developed a variety of generalization bounds for deep learning, all based on the fundamental learning-theoretic…
Let $A$ be a central division algebra of prime degree $p$ over $\mathbb{Q}$. We obtain subconvex hybrid bounds, uniform in both the eigenvalue and the discriminant, for the sup-norm of Hecke-Maass forms on the compact quotients of…
A priori, a posteriori, and mixed type upper bounds for the absolute change in Ritz values of self-adjoint matrices in terms of submajorization relations are obtained. Some of our results prove recent conjectures by Knyazev, Argentati, and…
We introduce several new notions of (sectional) curvature bounds for Lorentzian pre-length spaces: On the one hand, we provide convexity/concavity conditions for the (modified) time separation function, and, on the other hand, we study…
The aim of this paper is to deepen the convergence analysis of the scaled gradient projection (SGP) method, proposed by Bonettini et al. in a recent paper for constrained smooth optimization. The main feature of SGP is the presence of a…
A scaled conjugate gradient method that accelerates existing adaptive methods utilizing stochastic gradients is proposed for solving nonconvex optimization problems with deep neural networks. It is shown theoretically that, whether with…
The paper addresses parametric inequality systems described by polynomial functions in finite dimensions, where state-dependent infinite parameter sets are given by finitely many polynomial inequalities and equalities. Such systems can be…