Related papers: Sharp convergence bounds for sums of POD and SPOD …
Given a large set $U$ where each item $a\in U$ has weight $w(a)$, we want to estimate the total weight $W=\sum_{a\in U} w(a)$ to within factor of $1\pm\varepsilon$ with some constant probability $>1/2$. Since $n=|U|$ is large, we want to do…
We prove the Simons-Johnson theorem for the sums $S_n$ of $m$-dependent random variables, with exponential weights and limiting compound Poisson distribution $\CP(s,\lambda)$. More precisely, we give sufficient conditions for…
The main result of this paper is the following: for all $b \in \mathbb Z$ there exists $k=k(b)$ such that \[ \max \{ |A^{(k)}|, |(A+u)^{(k)}| \} \geq |A|^b, \] for any finite $A \subset \mathbb Q$ and any non-zero $u \in \mathbb Q$. Here,…
We establish the sharpness of the percolation phase transition for a class of infinite-range weighted random connection models. The vertex set is given by a marked Poisson point process on $\mathbb{R}^d$ with intensity $\lambda>0$, where…
A long time ago, it has been conjectured that a Hamiltonian with a potential of the form x^2+i v x^3, v real, has a real spectrum. This conjecture has been generalized to a class of so-called PT symmetric Hamiltonians and some proofs have…
We provide a simple algorithm for finding the optimal upper bound for sums of products of matrix entries of the form S_pi(N) := sum_{j_1, ..., j_2m = 1}^N t^1_{j_1 j_2} t^2_{j_3 j_4} ... t^m_{j_2m-1 j_2m} where some of the summation indices…
Stochastic gradient descent (SGD) and its variants are the main workhorses for solving large-scale optimization problems with nonconvex objective functions. Although the convergence of SGDs in the (strongly) convex case is well-understood,…
We prove an inequality on decision trees on monotonic measures which generalizes the OSSS inequality on product spaces. As an application, we use this inequality to prove a number of new results on lattice spin models and their…
We study population convergence guarantees of stochastic gradient descent (SGD) for smooth convex objectives in the interpolation regime, where the noise at optimum is zero or near zero. The behavior of the last iterate of SGD in this…
Under the assumption that the distribution of a nonnegative random variable $X$ admits a bounded coupling with its size biased version, we prove simple and strong concentration bounds. In particular the upper tail probability is shown to…
There is an extensive literature on the asymptotic order of Sudler's trigonometric product $P_N (\alpha) = \prod_{n=1}^N |2 \sin (\pi n \alpha)|$ for fixed or for "typical" values of $\alpha$. In the present paper we establish a structural…
We prove novel convergence results for a stochastic proximal gradient algorithm suitable for solving a large class of convex optimization problems, where a convex objective function is given by the sum of a smooth and a possibly non-smooth…
Let $U$ be a connected open subset of $\mathbb{R}^n$, and let $X=(X_1,X_{2},\ldots,X_m)$ be a system of H\"{o}rmander vector fields defined on $U$. This paper addresses sharp embedding results and geometric inequalities in the generalized…
Recent advances in randomized incremental methods for minimizing $L$-smooth $\mu$-strongly convex finite sums have culminated in tight complexity of $\tilde{O}((n+\sqrt{n L/\mu})\log(1/\epsilon))$ and $O(n+\sqrt{nL/\epsilon})$, where…
We consider partitions $p_{w}(n)$ of a positive integer $n$ arising from the generating functions \[ \sum_{n=1}^\infty p_{w}(n) z^n = \prod_{m \in \mathbb{N}} (1-z^m)^{-w(m)}, \] where the weights $w(m)$ are M\"{o}bius convolutions. We…
Using a one-to-one correspondence between observables and their spectral resolutions, we introduce the sum of any two bounded observables of a $\sigma$-MV-effect algebra. This sum is commutative, associative and with neutral element. Under…
Stochastic mirror descent (SMD) is a fairly new family of algorithms that has recently found a wide range of applications in optimization, machine learning, and control. It can be considered a generalization of the classical stochastic…
Let $\mathfrak{g}$ be a complex Kac-Moody algebra, with Cartan subalgebra $\mathfrak{h}$. Also fix a weight $\lambda\in\mathfrak{h}^*$. For $M(\lambda)\twoheadrightarrow V$ an arbitrary highest weight $\mathfrak{g}$-module, we provide a…
We consider the $\mathcal{N}=2$ SYM theory with gauge group SU($N$) and a matter content consisting of one multiplet in the symmetric and one in the anti-symmetric representation. This conformal theory admits a large-$N$ 't Hooft expansion…
Stochastic Gradient Descent (SGD) is a central tool in machine learning. We prove that SGD converges to zero loss, even with a fixed (non-vanishing) learning rate - in the special case of homogeneous linear classifiers with smooth monotone…