Related papers: A Note on Distribution Free Symmetrization Inequal…
We revisit the range sampling problem: the input is a set of points where each point is associated with a real-valued weight. The goal is to store them in a structure such that given a query range and an integer $k$, we can extract $k$…
In this paper we propose and study a class of nonparametric, yet interpretable measures of association between two random vectors $X$ and $Y$ taking values in $\mathbb{R}^{d_1}$ and $\mathbb{R}^{d_2}$ respectively ($d_1, d_2\ge 1$). These…
We prove that if $X$ and $Y$ are two Gaussian fields with equivalent spectral densities, they have the same sample paths properties in any separable Banach space continuously embedded in $\mathcal{C}^0(K)$ where $K$ is a compact set of…
Consider independent observations $(X_1,R_1)$, $(X_2,R_2)$, \ldots, $(X_n,R_n)$ with random or fixed ranks $R_i \in \{1,2,\ldots,k\}$, while conditional on $R_i = r$, the random variable $X_i$ has the same distribution as the $r$-th order…
This paper derives the asymptotic distribution of variance weighted Kolmogorov-Smirnov statistics for conditional moment inequality models for the case of a one dimensional covariate. The asymptotic distribution depends on the data…
We study the Banach space $D([0,1]^m)$ of functions of several variables that are (in a certain sense) right-continuous with left limits, and extend several results previously known for the standard case $m=1$. We give, for example, a…
High-dimensional k-sample comparison is a common applied problem. We construct a class of easy-to-implement nonparametric distribution-free tests based on new tools and unexplored connections with spectral graph theory. The test is shown to…
We develop novel empirical Bernstein inequalities for the variance of bounded random variables. Our inequalities hold under constant conditional variance and mean, without further assumptions like independence or identical distribution of…
In the setting where we have $n$ independent observations of a random variable $X$, we derive explicit error bounds in total variation distance when approximating the number of observations equal to the maximum of the sample (in the case…
The paper proposes one-to-one transformation of the vector of components $\{Y_{in}\}_{i=1}^m$ of Pearson's chi-square statistic, \[Y_{in}=\frac{\nu_{in}-np_i}{\sqrt{np_i}},\qquad i=1,\ldots,m,\] into another vector $\{Z_{in}\}_{i=1}^m$,…
We provide efficient algorithms for the problem of distribution learning from high-dimensional Gaussian data where in each sample, some of the variable values are missing. We suppose that the variables are missing not at random (MNAR). The…
We propose a framework for analyzing and comparing distributions, allowing us to design statistical tests to determine if two samples are drawn from different distributions. Our test statistic is the largest difference in expectations over…
We determine the joint limiting distribution of adjacent spacings around a central, intermediate, or an extreme order statistic $X_{k:n}$ of a random sample of size $n$ from a continuous distribution $F$. For central and intermediate cases,…
We show a deviation inequality for U-statistics of independent data taking values in a separable Banach space which satisfies some smoothness assumptions. We then provide applications to rates in the law of large numbers for U-statistics, a…
We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side…
We prove the following exponential inequality: Let $n\geq 1$ and let $X_1,...,X_n$ be $n$ independent identically distributed symmetric real-valued random variables. For any $x,y>0$, we have \[\mathbb{P}\big({X_1+...+X_n}\geq x,\,…
A basic result is that the sample variance for i.i.d. observations is an unbiased estimator of the variance of the underlying distribution (see for instance Casella and Berger (2002)). But what happens if the observations are neither…
Testing for association or dependence between pairs of random variables is a fundamental problem in statistics. In some applications, data are subject to selection bias that causes dependence between observations even when it is absent from…
We study finite-sample inference for the trade-off function of two unknown probability distributions, the function that traces the optimal type I/type II error frontier in binary testing. Given samples from distributions $P$ and $Q$, we…
In scientific studies involving analyses of multivariate data, basic but important questions often arise for the researcher: Is the sample exchangeable, meaning that the joint distribution of the sample is invariant to the ordering of the…