Related papers: A posteriori superlinear convergence bounds for bl…
This note provides a novel, simple analysis of the method of conjugate gradients for the minimization of convex quadratic functions. In contrast with standard arguments, our proof is entirely self-contained and does not rely on the…
Accelerated stochastic gradient descent (ASGD) is a workhorse in deep learning and often achieves better generalization performance than SGD. However, existing optimization theory can only explain the faster convergence of ASGD, but cannot…
Lower a posteriori error bounds obtained using the standard bubble function approach are reviewed in the context of anisotropic meshes. A numerical example is given that clearly demonstrates that the short-edge jump residual terms in such…
The Bounded Negativity Conjecture predicts that for any smooth complex surface $X$ there exists a lower bound for the selfintersection of reduced divisors on $X$. This conjecture is open. It is also not known if the existence of such a…
We prove the first inverse theorem for point--sphere incidence bounds over finite fields in dimensions $d \ge 3$, showing that near-extremality forces algebraic rigidity. While sharp upper bounds have been known for over a decade, the…
A while ago MLC (the conjecture that the Mandelbrot set is locally connected) was proven for quasi-hyperbolic points by Douady and Hubbard, and for boundaries of hyperbolic components by Yoccoz. More recently Yoccoz proved MLC for all at…
Motivated by problems on random differences in Szemer\'{e}di's theorem and on large deviations for arithmetic progressions in random sets, we prove upper bounds on the Gaussian width of point sets that are formed by the image of the…
In this paper, we present some explicit exponents in the estimates for the volumes of sub-level sets of polynomials on bounded sets, and applications to the decay of oscillatory integrals and the convergent of singular integrals.
The success of deep architectures is at least in part attributed to the layer-by-layer unsupervised pre-training that initializes the network. Various papers have reported extensive empirical analysis focusing on the design and…
A precise tie between a univariate spline's knots and its zeros abundance and dissemination is formulated. As an application, a conjecture formulated by De Concini and Procesi is shown to be true in the special univariate, unimodular case.…
We obtain upper bounds, independent of the ambient dimension, for the number of realizable zero-nonzero patterns and (over ordered fields) sign conditions of a finite family of polynomials $\mathcal P$ restricted to an algebraic subset $V$…
This paper is concerned with the convergence analysis of an extended variation of the locally optimal preconditioned conjugate gradient method (LOBPCG) for the extreme eigenvalue of a Hermitian matrix polynomial which admits some extended…
In this paper we propose a randomized primal-dual proximal block coordinate updating framework for a general multi-block convex optimization model with coupled objective function and linear constraints. Assuming mere convexity, we establish…
We study stochastic gradient descent (SGD) for composite optimization problems with $N$ sequential operators subject to perturbations in both the forward and backward passes. Unlike classical analyses that treat gradient noise as additive…
We present bounds for the geometric degree of the tangent bundle and the tangential variety of a smooth affine algebraic variety $V$ in terms of the geometric degree of $V$. We first analyze the case of curves, showing an explicit relation…
We prove the first superpolynomial lower bounds for learning one-layer neural networks with respect to the Gaussian distribution using gradient descent. We show that any classifier trained using gradient descent with respect to square-loss…
In the present work, we derive functional upper bounds for the potential error arising from finite-element boundary-element coupling formulations for a nonlinear Poisson-type transmission problem. The proposed a posteriori error estimates…
We study a general class of bilevel problems, consisting in the minimization of an upper-level objective which depends on the solution to a parametric fixed-point equation. Important instances arising in machine learning include…
We study widths of conjugacy classes in anisotropic higher rank $S$-arithmetic groups of orthogonal type. Assuming the GRH, we prove that many such groups have bounded conjugacy width. For example, this holds if the degree is greater or…
Two of the most popular parallel-in-time methods are Parareal and multigrid-reduction-in-time (MGRIT). Recently, a general convergence theory was developed in Southworth (2019) for linear two-level MGRIT/Parareal that provides necessary and…