Related papers: Lower tail large deviations of the stochastic six …
The paper suggests a simple method of deriving minimax lower bounds to the accuracy of statistical inference on heavy tails. A well-known result by Hall and Welsh (Ann. Statist. 12 (1984) 1079-1084) states that if $\hat{\alpha}_n$ is an…
We study stochastic convex optimization with heavy-tailed data under the constraint of differential privacy (DP). Most prior work on this problem is restricted to the case where the loss function is Lipschitz. Instead, as introduced by…
While the convergence behaviors of stochastic gradient methods are well understood \emph{in expectation}, there still exist many gaps in the understanding of their convergence with \emph{high probability}, where the convergence rate has a…
We study large deviation probabilities for a sum of dependent random variables from a heavy-tailed factor model, assuming that the components are regularly varying. We identify conditions where both the factor and the idiosyncratic terms…
We study large deviation properties of probability distributions with either a compact support or a fat tail by comparing them with q-deformed exponential distributions. Our main result is a large deviation property for probability…
Consider the upper tail probability that the homomorphism count of a fixed graph $H$ within a large sparse random graph $G_n$ exceeds its expected value by a fixed factor $1+\delta$. Going beyond the Erd\H{o}s-R\'enyi model, we establish…
In recent years, various notions of capacity and complexity have been proposed for characterizing the generalization properties of stochastic gradient descent (SGD) in deep learning. Some of the popular notions that correlate well with the…
Directed last passage percolation models on the plane, where one studies the weight as well as the geometry of optimizing paths (called polymers) in a field of i.i.d. weights, are paradigm examples of models in the KPZ universality class.…
We establish the large deviation probabilities for the height of random recursive trees, revealing polynomial upper-tail decay and stretched-exponential lower-tail decay. Remarkably, the lower tail features an atypical prefactor that grows…
For first passage percolation on $\mathbb{Z}^2$ with i.i.d. bounded edge weights, we consider the upper tail large deviation event; i.e., the rare situation where the first passage time between two points at distance $n$, is macroscopically…
Using tail bounds, we introduce a new probabilistic condition for function estimation in stochastic derivative-free optimization which leads to a reduction in the number of samples and eases algorithmic analyses. Moreover, we develop simple…
In recent works on the theory of machine learning, it has been observed that heavy tail properties of Stochastic Gradient Descent (SGD) can be studied in the probabilistic framework of stochastic recursions. In particular,…
This paper proposes a probabilistic approach to investigate the shape of landscapes of multi-dimensional potential functions. Under a suitable coupling scheme, two copies of the overdamped Langevin dynamics associated with the potential…
Estimating the probability of extreme events involving multiple risk factors is a critical challenge in fields such as finance and climate science. This paper proposes a semi-parametric approach to estimate the probability that a…
The study of loss function distributions is critical to characterize a model's behaviour on a given machine learning problem. For example, while the quality of a model is commonly determined by the average loss assessed on a testing set,…
In this work, we study the convergence \emph{in high probability} of clipped gradient methods when the noise distribution has heavy tails, ie., with bounded $p$th moments, for some $1<p\le2$. Prior works in this setting follow the same…
Convolutions of long-tailed and subexponential distributions play a major role in the analysis of many stochastic systems. We study these convolutions, proving some important new results through a simple and coherent approach, and showing…
We characterize the complex, heavy-tailed probability distribution functions (pdf) describing the response and its local extrema for structural systems subjected to random forcing that includes extreme events. Our approach is based on the…
We propose a distance supervised relation extraction approach for long-tailed, imbalanced data which is prevalent in real-world settings. Here, the challenge is to learn accurate "few-shot" models for classes existing at the tail of the…
Starting with the large deviation principle (LDP) for the Erd\H{o}s-R\'enyi binomial random graph $\mathcal{G}(n,p)$ (edge indicators are i.i.d.), due to Chatterjee and Varadhan (2011), we derive the LDP for the uniform random graph…