English
Related papers

Related papers: On Identity Tests for High Dimensional Data Using …

200 papers

Asymptotic methods for hypothesis testing in high-dimensional data usually require the dimension of the observations to increase to infinity, often with an additional condition on its rate of increase compared to the sample size. On the…

Statistics Theory · Mathematics 2024-03-26 Joydeep Chowdhury , Subhajit Dutta , Marc G. Genton

High dimensional data analysis for exploration and discovery includes three fundamental tasks: dimensionality reduction, clustering, and visualization. When the three associated tasks are done separately, as is often the case thus far,…

Machine Learning · Computer Science 2020-12-02 Stan Z. Li , Lirong Wu , Zelin Zang

A new family of nonparametric statistics, the r-statistics, is introduced. It consists of counting the number of records of the cumulative sum of the sample. The single-sample r-statistic is almost as powerful as Student's t-statistic for…

Methodology · Statistics 2015-07-14 Damien Challet

Two-sample testing is a fundamental problem in statistics. Despite its long history, there has been renewed interest in this problem with the advent of high-dimensional and complex data. Specifically, in the machine learning literature,…

Methodology · Statistics 2019-11-19 Ilmun Kim , Ann B. Lee , Jing Lei

Testing the equality in distributions of multiple samples is a common task in many fields. However, this problem for high-dimensional or non-Euclidean data has not been well explored. In this paper, we propose new nonparametric tests based…

Methodology · Statistics 2022-05-30 Hoseung Song , Hao Chen

We consider the problem of testing whether a correlation matrix of a multivariate normal population is the identity matrix. We focus on sparse classes of alternatives where only a few entries are nonzero and, in fact, positive. We derive a…

Statistics Theory · Mathematics 2015-04-15 Ery Arias-Castro , Sébastien Bubeck , Gábor Lugosi

In this paper, we develop a systematic theory for high dimensional analysis of variance in multivariate linear regression, where the dimension and the number of coefficients can both grow with the sample size. We propose a new \emph{U}~type…

Methodology · Statistics 2023-01-12 Zhipeng Lou , Xianyang Zhang , Wei Biao Wu

This article introduces a robust hypothesis testing procedure: the Lq-likelihood-ratio-type test (LqRT). By deriving the asymptotic distribution of this test statistic, the authors demonstrate its robustness both analytically and…

Applications · Statistics 2016-09-27 Yichen Qin , Carey E. Priebe

For $k,m,n\in \mathbb{N}$, we consider $n^k\times n^k$ random matrices of the form $$ \mathcal{M}_{n,m,k}(\mathbf{y})=\sum_{\alpha=1}^m\tau_\alpha {Y_\alpha}Y_\alpha^T,\quad…

Probability · Mathematics 2017-01-27 Anna Lytova

Often when we deal with `Big Data', the true effects we are interested in are Rare and Weak (RW). Researchers measure a large number of features, hoping to find perhaps only a small fraction of them to be relevant to the research in…

Statistics Theory · Mathematics 2014-10-20 Jiashun Jin , Tracy Ke

We consider two hypothesis testing problems for low-rank and high-dimensional tensor signals, namely the tensor signal alignment and tensor signal matching problems. These problems are challenging due to the high dimension of tensors and…

Methodology · Statistics 2026-02-10 Ruihan Liu , Zhenggang Wang , Jianfeng Yao

Nonparametric two sample testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. We refer to the most common…

Statistics Theory · Mathematics 2015-08-05 Aaditya Ramdas , Sashank J. Reddi , Barnabas Poczos , Aarti Singh , Larry Wasserman

LLM-generated tabular data is creating new opportunities for data-driven applications in academia, business, and society. To leverage benefits like missing value imputation, labeling, and enrichment with context-aware attributes,…

Human-Computer Interaction · Computer Science 2025-05-08 Madhav Sachdeva , Christopher Narayanan , Marvin Wiedenkeller , Jana Sedlakova , Jürgen Bernard

In this paper we propose a new approach to the central limit theorem (CLT), based on functions of bounded F\'echet variation for the continuously differentiable linear statistics of random matrix ensembles which relies on: a weaker form of…

Probability · Mathematics 2022-01-12 Mario Diaz , James A. Mingo

Clustering methods have led to a number of important discoveries in bioinformatics and beyond. A major challenge in their use is determining which clusters represent important underlying structure, as opposed to spurious sampling artifacts.…

Methodology · Statistics 2021-10-20 Hanwen Huang , Yufeng Liu , Ming Yuan , J. S. Marron

This work is motivated by learning the individualized minimal clinically important difference, a vital concept to assess clinical importance in various biomedical studies. We formulate the scientific question into a high-dimensional…

Methodology · Statistics 2023-03-28 Huijie Feng , Jingyi Duan , Yang Ning , Jiwei Zhao

Consider a $N\times n$ matrix $\Sigma_n=\frac{1}{\sqrt{n}}R_n^{1/2}X_n$, where $R_n$ is a nonnegative definite Hermitian matrix and $X_n$ is a random matrix with i.i.d. real or complex standardized entries. The fluctuations of the linear…

Probability · Mathematics 2016-06-29 Jamal Najim , Jianfeng Yao

Nonparametric tests for functional data are a challenging class of tests to work with because of the potentially high dimensional nature of the data. One of the main challenges for considering rank-based tests, like the Mann-Whitney or…

Methodology · Statistics 2024-07-12 Mark J. Meyer

Two new symmetry tests, of integral and Kolmogorov type, based on the characterization by squares of linear statistics are proposed. The test statistics are related to the family of degenerate U-statistics. Their asymptotic properties are…

Methodology · Statistics 2023-05-30 V. Božin , B. Milošević , Ya. Yu. Nikitin , M. Obradović

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

Statistics Theory · Mathematics 2020-03-09 Rémy Mariétan , Stephan Morgenthaler