Related papers: Time-Uniform Self-Normalized Concentration for Vec…
Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial…
This paper derives non-asymptotic error bounds for nonlinear stochastic approximation algorithms in the Wasserstein-$p$ distance. To obtain explicit finite-sample guarantees for the last iterate, we develop a coupling argument that compares…
Due to their flexibility, Gaussian processes (GPs) have been widely used in nonparametric function estimation. A prior information about the underlying function is often available. For instance, the physical system (computer model output)…
We modify the classical Bernstein's inequality for the sums of independent centered random variables (r.v.) in the terms of relative tails or moments. We built also some examples in order to show the exactness of offered results.
The seminal papers of Pickands [1,2] paved the way for a systematic study of high exceedance probabilities of both stationary and non-stationary Gaussian processes. Yet, in the vector-valued setting, due to the lack of key tools including…
Training large neural networks exposes neural scaling laws for the generalization error, which points to a universal behavior across network architectures of learning in high dimensions. It was also shown that this effect persists in the…
This paper gives two theoretical results on estimating low-rank parameter matrices for linear models with multivariate responses. We first focus on robust parameter estimation of low-rank multi-task learning with heavy-tailed data and…
We investigate extreme value theory of a class of random sequences defined by the all-time suprema of aggregated self-similar Gaussian processes with trend. This study is motivated by its potential applications in various areas and its…
We derive, up to a constant factor, matching lower and upper bounds on the concentration functions of suprema of separable centered Gaussian processes and order statistics of Gaussian random fields. These bounds reveal that suprema of…
Let $\{X, X_n, n\geq 1\}$ be a sequence of independent identically distributed non-degenerate random variables. Put $S_0=0, S_n = \sum^n_{i=1} X_i$ and $V_n^2=\sum^n_{i=1} X_i^2, n\ge 1.$ A weak convergence theorem is established for the…
Extreme-value theory for random vectors and stochastic processes with continuous trajectories is usually formulated for random objects all of whose univariate marginal distributions are identical. In the spirit of Sklar's theorem from…
Consider $n$ i.i.d. random elements on $C[0,1]$. We show that, under an appropriate strengthening of the domain of attraction condition, natural estimators of the extreme-value index, which is now a continuous function, and the normalizing…
Following the student t-statistic, normalization has been a widely used method in statistic and other disciplines including economics, ecology and machine learning. We focus on statistics taking the form of a ratio over (some power of) the…
We are concerned with obtaining novel concentration inequalities for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We not only derive - for the first time - distribution-free Bernstein-like…
Motivated by Talagrand's conjecture on regularization properties of the natural semigroup on the Boolean hypercube, and in particular its continuous analogue involving regularization properties of the Ornstein-Uhlenbeck semigroup acting on…
This work is concerned with the convergence of Gaussian process regression. A particular focus is on hierarchical Gaussian process regression, where hyper-parameters appearing in the mean and covariance structure of the Gaussian process…
Following the concentration of the measure theory formalism, we consider the transformation $\Phi(Z)$ of a random variable $Z$ having a general concentration function $\alpha$. If the transformation $\Phi$ is $\lambda$-Lipschitz with…
Recent advances in semi-supervised learning have shown tremendous potential in overcoming a major barrier to the success of modern machine learning algorithms: access to vast amounts of human-labeled training data. Previous algorithms based…
We consider a positive stationary generalized Ornstein--Uhlenbeck process \[V_t=\mathrm{e}^{-\xi_t}\biggl(\int_0^t\mathrm{e}^{\xi_{s-}}\ ,\mathrm{d}\eta_s+V_0\biggr)\qquadfor t\geq0,\] and the increments of the integrated generalized…
Many sequence-to-sequence tasks in natural language processing are roughly monotonic in the alignment between source and target sequence, and previous work has facilitated or enforced learning of monotonic attention behavior via specialized…