Related papers: Lower tail large deviations of the stochastic six …
For large model spaces, the potential entrapment of Markov chain Monte Carlo (MCMC) based methods with spike-and-slab priors poses significant challenges in posterior computation in regression models. On the other hand, maximum a posteriori…
Using terminologies of information geometry, we derive upper and lower bounds of the tail probability of the sample mean. Employing these bounds, we obtain upper and lower bounds of the minimum error probability of the 2nd kind of error…
We present the derivation of a new model to describe neutron spin echo spectroscopy and quasi-elastic neutron scattering data on liposomes. We compare the new model with existing approaches and benchmark it with experimental data. The…
Stochastic gradient descent (SGD) with mini-batching is a standard tool in large-scale optimization, yet its theoretical properties under heavy-tailed gradient noise remain largely unexplored. In this paper we study SGD with increasing…
We give a sufficient condition for the exponential decay of the tail probability of a non-negative random variable. We consider the Laplace-Stieltjes transform of the probability distribution function of the random variable. We present a…
Adaptive optimization methods (such as Adam) play a major role in LLM pretraining, significantly outperforming Gradient Descent (GD). Recent studies have proposed new smoothness assumptions on the loss function to explain the advantages of…
Extreme events and the heavy tail distributions driven by them are ubiquitous in various scientific, engineering and financial research. They are typically associated with stochastic instability caused by hidden unresolved processes.…
Let $\Psi_1,\Psi_2,...$ be a sequence of i.i.d. random Lipschitz functions on a complete separable metric space with unbounded metric $d$ and forward iterations $X_n$. Suppose that $X_n$ has a stationary distribution. We study the…
We consider a regression framework where the design points are deterministic and the errors possibly non-i.i.d. and heavy-tailed (with a moment of order $p$ in $[1,2]$). Given a class of candidate regression functions, we propose a…
In this paper we propagate a large deviations approach for proving limit theory for (generally) multivariate time series with heavy tails. We make this notion precise by introducing regularly varying time series. We provide general large…
Given a branching random walk $(Z_n)_{n\geq0}$ on $\mathbb{R}$, let $Z_n(A)$ be the number of particles located in interval $A$ at generation $n$. It is well known (e.g., \cite{biggins}) that under some mild conditions, $Z_n(\sqrt…
Neural network compression techniques have become increasingly popular as they can drastically reduce the storage and computation requirements for very large networks. Recent empirical studies have illustrated that even simple pruning…
Consider the short-time probability distribution $\mathcal{P}(H,t)$ of the one-point interface height difference $h(x=0,\tau=t)-h(x=0,\tau=0)=H$ of the stationary interface $h(x,\tau)$ described by the Kardar-Parisi-Zhang equation. It was…
Stochastic first-order methods such as Stochastic Extragradient (SEG) or Stochastic Gradient Descent-Ascent (SGDA) for solving smooth minimax problems and, more generally, variational inequality problems (VIP) have been gaining a lot of…
Quantifying tail dependence is an important issue in insurance and risk management. The prevalent tail dependence coefficient (TDC), however, is known to underestimate the degree of tail dependence and it does not capture non-exchangeable…
We study in this paper the problem of least absolute deviation (LAD) regression for high-dimensional heavy-tailed time series which have finite $\alpha$-th moment with $\alpha \in (1,2]$. To handle the heavy-tailed dependent data, we…
Despite the successes of probabilistic models based on passing noise through neural networks, recent work has identified that such methods often fail to capture tail behavior accurately, unless the tails of the base distribution are…
Recent studies have shown that gradient descent (GD) can achieve improved generalization when its dynamics exhibits a chaotic behavior. However, to obtain the desired effect, the step-size should be chosen sufficiently large, a task which…
This paper is concerned with the general theme of relating the Large Deviation Principle (LDP) for the invariant measures of stochastic processes to the associated sample path LDP. It is shown that if the sample path deviation function…
Stochastic volatility processes with heavy-tailed innovations are a well-known model for financial time series. In these models, the extremes of the log returns are mainly driven by the extremes of the i.i.d. innovation sequence which leads…