Related papers: Symmetric KL-divergence by Stein's Method
Recently, many machine learning and statistical models such as non-linear regressions, the Single Index, Multi-index, Varying Coefficient Index Models and Two-layer Neural Networks can be reduced to or be seen as a special case of a new…
In this paper, we consider the problem of Gaussian approximation for the online linear regression task. We derive the corresponding rates for the setting of a constant learning rate and study the explicit dependence of the convergence rate…
We consider distributed optimization on undirected connected graphs. We propose a novel distributed conditional gradient method with (O(1/\sqrt{k})) convergence. Compared with existing methods, each iteration of our method uses both…
We show how to determine whether a given pattern p of length m occurs in a given text t of length n in ${\tilde O}(\sqrt{n}+\sqrt{m})$\footnote{${\tilde O}$ allows for logarithmic factors in m and $n/m$} time, with inverse polynomial…
Let $\boldsymbol{\xi}=(\xi_1,\ldots,\xi_m)$ be a negatively associated mean zero random vector with components that obey the bound $|\xi_i| \le B, i=1,\ldots,m$, and whose sum $W = \sum_{i=1}^m \xi_i$ has variance 1, the bound \[…
By using the pseudo-metric introduced in [F. Golse, T. Paul: Archive for Rational Mech. Anal. 223 (2017) 57-94], which is an analogue of the Wasserstein distance of exponent $2$ between a quantum density operator and a classical…
Let $\varphi_{n,K}$ denote the largest angle in all the triangles with vertices among the $n$ points selected at random in a compact convex subset $K$ of $\mathbb{R}^d$ with nonempty interior, where $d\ge2$. It is shown that the…
By a delicate analysis for the Stein's equation associated to the $\alpha$-stable law approximation with $\alpha \in (0,2)$, we prove a quantitative stable central limit theorem in Wasserstein type distance, which generalizes the results in…
In this paper, we present a minimal formalism for Stein operators which leads to different probabilistic representations of solutions to Stein equations. These in turn provide a wide family of Stein-Covariance identities which we put to use…
The problem of estimating the Kullback-Leibler divergence $D(P\|Q)$ between two unknown distributions $P$ and $Q$ is studied, under the assumption that the alphabet size $k$ of the distributions can scale to infinity. The estimation is…
Elkies and McMullen [Duke Math.J.~123 (2004) 95--139] have shown that the gaps between the fractional parts of \sqrt n for n=1,\ldots,N, have a limit distribution as N tends to infinity. The limit distribution is non-standard and differs…
We present a mathematical analysis of the Wang-Landau algorithm, prove its convergence, identify sources of errors and strategies for optimization. In particular, we found the histogram increases uniformly with small fluctuation after a…
The discrete distribution of the length of longest increasing subsequences in random permutations of $n$ integers is deeply related to random matrix theory. In a seminal work, Baik, Deift and Johansson provided an asymptotics in terms of…
We consider the extreme value statistics of correlated random variables that arise from a Langevin equation. Recently, it was shown that the extreme values of the Ornstein-Uhlenbeck process follow a different distribution than those…
We provide non-asymptotic convergence rates of the Polyak-Ruppert averaged stochastic gradient descent (SGD) to a normal random vector for a class of twice-differentiable test functions. A crucial intermediate step is proving a…
The $k$-Opt and Lin-Kernighan algorithm are two of the most important local search approaches for the Metric TSP. Both start with an arbitrary tour and make local improvements in each step to get a shorter tour. We show that for any fixed…
In this paper, we quantitative convergence in $W_2$ for a family of Langevin-like stochastic processes that includes stochastic gradient descent and related gradient-based algorithms. Under certain regularity assumptions, we show that the…
We establish Cram\'er-type moderate deviation theorems for sums of locally dependent random variables and combinatorial central limit theorems. Under some mild exponential moment conditions, optimal error bounds and convergence ranges are…
Comparing probability distributions is an indispensable and ubiquitous task in machine learning and statistics. The most common way to compare a pair of Borel probability measures is to compute a metric between them, and by far the most…
Stein's method for Gaussian process approximation can be used to bound the differences between the expectations of smooth functionals $h$ of a c\`adl\`ag random process $X$ of interest and the expectations of the same functionals of a well…