English
Related papers

Related papers: Time-Uniform Self-Normalized Concentration for Vec…

200 papers

Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial…

Computation and Language · Computer Science 2018-06-19 Matthias Sperber , Jan Niehues , Graham Neubig , Sebastian Stüker , Alex Waibel

This paper derives non-asymptotic error bounds for nonlinear stochastic approximation algorithms in the Wasserstein-$p$ distance. To obtain explicit finite-sample guarantees for the last iterate, we develop a coupling argument that compares…

Machine Learning · Computer Science 2026-02-03 Seo Taek Kong , R. Srikant

Due to their flexibility, Gaussian processes (GPs) have been widely used in nonparametric function estimation. A prior information about the underlying function is often available. For instance, the physical system (computer model output)…

Methodology · Statistics 2017-11-21 Hassan Maatouk

We modify the classical Bernstein's inequality for the sums of independent centered random variables (r.v.) in the terms of relative tails or moments. We built also some examples in order to show the exactness of offered results.

Probability · Mathematics 2022-06-03 M. R. Formica , E. Ostrovsky , L. Sirota

The seminal papers of Pickands [1,2] paved the way for a systematic study of high exceedance probabilities of both stationary and non-stationary Gaussian processes. Yet, in the vector-valued setting, due to the lack of key tools including…

Probability · Mathematics 2019-11-18 Krzysztof Dȩbicki , Enkelejd Hashorva , Longmin Wang

Training large neural networks exposes neural scaling laws for the generalization error, which points to a universal behavior across network architectures of learning in high dimensions. It was also shown that this effect persists in the…

Disordered Systems and Neural Networks · Physics 2026-02-27 Jakob Kramp , Javed Lindner , Moritz Helias

This paper gives two theoretical results on estimating low-rank parameter matrices for linear models with multivariate responses. We first focus on robust parameter estimation of low-rank multi-task learning with heavy-tailed data and…

Statistics Theory · Mathematics 2023-05-24 Kangqiang Li , Yuxuan Wang

We investigate extreme value theory of a class of random sequences defined by the all-time suprema of aggregated self-similar Gaussian processes with trend. This study is motivated by its potential applications in various areas and its…

Probability · Mathematics 2022-11-09 Lanpeng Ji , Xiaofan Peng

We derive, up to a constant factor, matching lower and upper bounds on the concentration functions of suprema of separable centered Gaussian processes and order statistics of Gaussian random fields. These bounds reveal that suprema of…

Probability · Mathematics 2023-10-19 Alexander Giessing

Let $\{X, X_n, n\geq 1\}$ be a sequence of independent identically distributed non-degenerate random variables. Put $S_0=0, S_n = \sum^n_{i=1} X_i$ and $V_n^2=\sum^n_{i=1} X_i^2, n\ge 1.$ A weak convergence theorem is established for the…

Probability · Mathematics 2013-06-21 Miklós Csörgő , Zhishui Hu

Extreme-value theory for random vectors and stochastic processes with continuous trajectories is usually formulated for random objects all of whose univariate marginal distributions are identical. In the spirit of Sklar's theorem from…

Probability · Mathematics 2016-12-23 Anne Sabourin , Johan Segers

Consider $n$ i.i.d. random elements on $C[0,1]$. We show that, under an appropriate strengthening of the domain of attraction condition, natural estimators of the extreme-value index, which is now a continuous function, and the normalizing…

Statistics Theory · Mathematics 2007-06-13 John H. J. Einmahl , Tao Lin

Following the student t-statistic, normalization has been a widely used method in statistic and other disciplines including economics, ecology and machine learning. We focus on statistics taking the form of a ratio over (some power of) the…

Statistics Theory · Mathematics 2025-09-19 Haolin Zou , Heyuan Yao , Victor de la Peña

We are concerned with obtaining novel concentration inequalities for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We not only derive - for the first time - distribution-free Bernstein-like…

Machine Learning · Statistics 2015-06-22 Bahman Yari Saeed Khanloo , Gholamreza Haffari

Motivated by Talagrand's conjecture on regularization properties of the natural semigroup on the Boolean hypercube, and in particular its continuous analogue involving regularization properties of the Ornstein-Uhlenbeck semigroup acting on…

Probability · Mathematics 2020-02-07 Nathael Gozlan , Mokshay Madiman , Cyril Roberto , Paul-Marie Samson

This work is concerned with the convergence of Gaussian process regression. A particular focus is on hierarchical Gaussian process regression, where hyper-parameters appearing in the mean and covariance structure of the Gaussian process…

Numerical Analysis · Mathematics 2020-07-20 Aretha L Teckentrup

Following the concentration of the measure theory formalism, we consider the transformation $\Phi(Z)$ of a random variable $Z$ having a general concentration function $\alpha$. If the transformation $\Phi$ is $\lambda$-Lipschitz with…

Probability · Mathematics 2026-02-03 Cosme Louart

Recent advances in semi-supervised learning have shown tremendous potential in overcoming a major barrier to the success of modern machine learning algorithms: access to vast amounts of human-labeled training data. Previous algorithms based…

Machine Learning · Computer Science 2019-11-22 Phi Vu Tran

We consider a positive stationary generalized Ornstein--Uhlenbeck process \[V_t=\mathrm{e}^{-\xi_t}\biggl(\int_0^t\mathrm{e}^{\xi_{s-}}\ ,\mathrm{d}\eta_s+V_0\biggr)\qquadfor t\geq0,\] and the increments of the integrated generalized…

Statistics Theory · Mathematics 2010-02-24 Vicky Fasen

Many sequence-to-sequence tasks in natural language processing are roughly monotonic in the alignment between source and target sequence, and previous work has facilitated or enforced learning of monotonic attention behavior via specialized…

Computation and Language · Computer Science 2021-04-09 Annette Rios , Chantal Amrhein , Noëmi Aepli , Rico Sennrich
‹ Prev 1 8 9 10 Next ›