English
Related papers

Related papers: A Bernstein-type inequality for suprema of random …

200 papers

This paper considers the problem of regression over distributions, which is becoming increasingly important in machine learning. Existing approaches often ignore the geometry of the probability space or are computationally expensive. To…

Machine Learning · Computer Science 2025-10-31 Maksim Maslov , Alexander Kugaevskikh , Matthew Ivanov

We propose a new estimator for the high-dimensional linear regression model with observation error in the design where the number of coefficients is potentially larger than the sample size. The main novelty of our procedure is that the…

Methodology · Statistics 2019-09-09 Alexandre Belloni , Abhishek Kaul , Mathieu Rosenbaum

Abstract In Extreme Value methodology the choice of threshold plays an important role in efficient modelling of observations exceeding the threshold. The threshold must be chosen high enough to ensure an unbiased extreme value index but…

Methodology · Statistics 2020-06-11 Andréhette Verster , Lizanne Raubenheimer

Concentration inequalities form an essential toolkit in the study of high dimensional (HD) statistical methods. Most of the relevant statistics literature in this regard is based on sub-Gaussian or sub-exponential tail assumptions. In this…

Statistics Theory · Mathematics 2023-01-09 Arun Kumar Kuchibhotla , Abhishek Chakrabortty

We show sharp bounds for probabilities of large deviations for sums of independent random variables satisfying Bernstein's condition. One such bound is very close to the tail of the standard Gaussian law in certain case; other bounds…

Probability · Mathematics 2015-07-13 Xiequan Fan , Ion Grama , Quansheng Liu

We study the sample complexity of estimating the covariance matrix $T$ of a distribution $\mathcal{D}$ over $d$-dimensional vectors, under the assumption that $T$ is Toeplitz. This assumption arises in many signal processing problems, where…

Signal Processing · Electrical Eng. & Systems 2019-10-31 Yonina C. Eldar , Jerry Li , Cameron Musco , Christopher Musco

We introduce a novel Bayesian estimator for the class proportion in an unlabeled dataset, based on the targeted learning framework. Our procedure requires the specification of a prior (and outputs a posterior) only for the target of…

Methodology · Statistics 2019-11-26 Iván Díaz , Oleksander Savenkov , Hooman Kamel

We develop a class of exponential bounds for the probability that a martingale sequence crosses a time-dependent linear threshold. Our key insight is that it is both natural and fruitful to formulate exponential concentration inequalities…

Probability · Mathematics 2025-12-18 Steven R. Howard , Aaditya Ramdas , Jon McAuliffe , Jasjeet Sekhon

We derive concentration inequalities for empirical means $\frac{1}{t} \int_0^t f(X_s) ds$ where $X_s$ is an irreducible Markov jump process on a finite state space and $f$ some observable. Using a Feynman-Kac semigroup we first derive a…

Probability · Mathematics 2022-10-13 Santiago Carrero Ibanez

In this paper, we consider the problem of column subset selection. We present a novel analysis of the spectral norm reconstruction for a simple randomized algorithm and establish a new bound that depends explicitly on the sampling…

Numerical Analysis · Mathematics 2015-05-05 Tianbao Yang , Lijun Zhang , Rong Jin , Shenghuo Zhu

We consider the problem of discriminative factor analysis for data that are in general non-Gaussian. A Bayesian model based on the ranks of the data is proposed. We first introduce a new {\em max-margin} version of the rank-likelihood. A…

Machine Learning · Statistics 2015-05-20 Xin Yuan , Ricardo Henao , Ephraim L. Tsalik , Raymond J. Langley , Lawrence Carin

The maximum-entropy remote sampling problem (MERSP) is to select a subset of s random variables from a set of n random variables, so as to maximize the information concerning a set of target random variables that are not directly…

Optimization and Control · Mathematics 2026-02-03 Gabriel Ponte , Marcia Fampa , Jon Lee

We present large sample results for partitioning-based least squares nonparametric regression, a popular method for approximating conditional expectation functions in statistics, econometrics, and machine learning. First, we obtain a…

Statistics Theory · Mathematics 2020-07-20 Matias D. Cattaneo , Max H. Farrell , Yingjie Feng

The seminal papers of Pickands [1,2] paved the way for a systematic study of high exceedance probabilities of both stationary and non-stationary Gaussian processes. Yet, in the vector-valued setting, due to the lack of key tools including…

Probability · Mathematics 2019-11-18 Krzysztof Dȩbicki , Enkelejd Hashorva , Longmin Wang

We consider the problem of estimating the mean $f$ of a Gaussian vector $Y$ with independent components of common unknown variance $\sigma^{2}$. Our estimation procedure is based on estimator selection. More precisely, we start with an…

Statistics Theory · Mathematics 2011-06-24 Yannick Baraud , Christophe Giraud , Sylvie Huet

We study sample average approximations (SAA) of chance constrained programs. SAA methods typically approximate the actual distribution in the chance constraint using an empirical distribution constructed from random samples assumed to be…

Optimization and Control · Mathematics 2022-05-13 Shuhao Yan , Francesca Parise , Eilyan Bitar

The surface area of a set which is only observed as a binary pixel image is often estimated by a weighted sum of pixel configurations counts. In this paper we examine these estimators in a design based setting -- we assume that the observed…

Statistics Theory · Mathematics 2019-06-20 Jürgen Kampf

This paper considers the entropy of the sum of (possibly dependent and non-identically distributed) Bernoulli random variables. Upper bounds on the error that follows from an approximation of this entropy by the entropy of a Poisson random…

Information Theory · Computer Science 2016-11-17 Igal Sason

In this paper, we are concerned with obtaining distribution-free concentration inequalities for mixture of independent Bernoulli variables that incorporate a notion of variance. Missing mass is the total probability mass associated to the…

Machine Learning · Statistics 2015-03-05 Bahman Yari Saeed Khanloo

In the study of the supremum of stochastic processes, Talagrand's chaining functionals and his generic chaining method are heavily related to the distribution of stochastic processes. In the present paper, we construct Talagrand's type…

Probability · Mathematics 2023-09-12 Yiming Chen , Pengtao Li , Dali Liu , Hanchao Wang