Related papers: On Identity Tests for High Dimensional Data Using …
We show that the variance of centred linear statistics of eigenvalues of GUE matrices remains bounded for large $n$ for some classes of test functions less regular than Lipschitz functions. This observation is suggested by the limiting form…
Recent advances have shown that statistical tests for the rank of cross-covariance matrices play an important role in causal discovery. These rank tests include partial correlation tests as special cases and provide further graphical…
Despite many applications, dimensionality reduction in the $\ell_1$-norm is much less understood than in the Euclidean norm. We give two new oblivious dimensionality reduction techniques for the $\ell_1$-norm which improve exponentially…
Two-sample tests for multivariate data and especially for non-Euclidean data are not well explored. This paper presents a novel test statistic based on a similarity graph constructed on the pooled observations from the two samples. It can…
This paper focuses on the prominent sphericity test when the dimension $p$ is much lager than sample size $n$. The classical likelihood ratio test(LRT) is no longer applicable when $p\gg n$. Therefore a Quasi-LRT is proposed and asymptotic…
Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…
We study sample covariance matrices of the form $W=\frac 1n C C^T$, where $C$ is a $k\times n$ matrix with i.i.d. mean zero entries. This is a generalization of so-called Wishart matrices, where the entries of $C$ are independent and…
We study the fluctuations of the eigenvalues of real valued large centrosymmetric random matrices via its linear eigenvalue statistic. This is essentially a central limit theorem (CLT) for sums of dependent random variables. The dependence…
Inference based on the penalized density ratio model is proposed and studied. The model under consideration is specified by assuming that the log--likelihood function of two unknown densities is of some parametric form. The model has been…
We address the issue of performing testing inference in generalized linear models when the sample size is small. This class of models provides a straightforward way of modeling normal and non-normal data and has been widely used in several…
Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessment. Currently, when such statistical measures are reported,…
This paper is concerned with Spearman's correlation matrices under large dimensional regime, in which the data dimension diverges to infinity proportionally with the sample size. We establish the central limit theorem for the linear…
We propose a class of nonparametric two-sample tests with a cost linear in the sample size. Two tests are given, both based on an ensemble of distances between analytic functions representing each of the distributions. The first test uses…
Consider a $N\times n$ random matrix $Y_n=(Y_{ij}^{n})$ where the entries are given by $$ Y_{ij}^{n}=\frac{\sigma_{ij}(n)}{\sqrt{n}} X_{ij}^{n} $$ the $X_{ij}^{n}$ being centered, independent and identically distributed random variables…
We propose a test of many zero parameter restrictions in a high dimensional linear iid regression model with $k$ $>>$ $n$ regressors. The test statistic is formed by estimating key parameters one at a time based on many low dimension…
Statistical significance testing is widely accepted as a means to assess how well a difference in effectiveness reflects an actual difference between systems, as opposed to random noise because of the selection of topics. According to…
Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement. But new biases are introduced due to the same LLM acting as both…
Testing whether the observed data conforms to a purported model (probability distribution) is a basic and fundamental statistical task, and one that is by now well understood. However, the standard formulation, identity testing, fails to…
The statistical analysis of discrete data has been the subject of extensive statistical research dating back to the work of Pearson. In this survey we review some recently developed methods for testing hypotheses about high-dimensional…
Asymptotic methods for hypothesis testing in high-dimensional data usually require the dimension of the observations to increase to infinity, often with an additional relationship between the dimension (say, $p$) and the sample size (say,…