English
Related papers

Related papers: Distribution-free Detection of a Submatrix

200 papers

Given a large matrix containing independent data entries, we consider the problem of detecting a submatrix inside the data matrix that contains larger-than-usual values. Different from previous literature, we do not have exact information…

Statistics Theory · Mathematics 2020-03-03 Yuchao Liu , Jiaqi Guo

The scan statistic is by far the most popular method for anomaly detection, being popular in syndromic surveillance, signal and image processing, and target detection based on sensor networks, among other applications. The use of the scan…

Methodology · Statistics 2016-11-28 Ery Arias-Castro , Rui M. Castro , Ervin Tánczos , Meng Wang

We consider the problem of detecting an elevated mean on an interval with unknown location and length in the univariate Gaussian sequence model. Recent results have shown that using scale-dependent critical values for the scan statistic…

Statistics Theory · Mathematics 2021-07-20 Guenther Walther , Andrew Perry

We consider the problem of localizing a submatrix with larger-than-usual entry values inside a data matrix, without the prior knowledge of the submatrix size. We establish an optimization framework based on a multiscale scan statistic, and…

Statistics Theory · Mathematics 2019-06-24 Yuchao Liu , Ery Arias-Castro

Anomaly detection when observing a large number of data streams is essential in a variety of applications, ranging from epidemiological studies to monitoring of complex systems. High-dimensional scenarios are usually tackled with…

Methodology · Statistics 2025-12-18 Ivo V. Stoepker , Rui M. Castro , Ery Arias-Castro , Edwin van den Heuvel

Classical two-sample permutation tests for equality of distributions have exact size in finite samples, but they fail to control size for testing equality of parameters that summarize each distribution. This paper proposes permutation tests…

Econometrics · Economics 2022-04-22 Marinho Bertanha , EunYi Chung

Datasets from the fields of bioinformatics, chemometrics, and face recognition are typically characterized by small samples of high-dimensional data. Among the many variants of linear discriminant analysis that have been proposed in order…

Machine Learning · Statistics 2020-04-20 Lama B. Niyazi , Abla Kammoun , Hayssam Dahrouj , Mohamed-Slim Alouini , Tareq Y. Al-Naffouri

We study the distributional properties of the linear discriminant function under the assumption of normality by comparing two groups with the same covariance matrix but different mean vectors. A stochastic representation for the…

Statistics Theory · Mathematics 2017-05-09 Taras Bodnar , Stepan Mazur , Edward Ngailo , Nestor Parolya

We consider large non-Hermitian random matrices $X$ with complex, independent, identically distributed centred entries and show that the linear statistics of their eigenvalues are asymptotically Gaussian for test functions having…

Probability · Mathematics 2023-10-16 Giorgio Cipolloni , László Erdős , Dominik Schröder

Rapid progress in representation learning has led to a proliferation of embedding models, and to associated challenges of model selection and practical application. It is non-trivial to assess a model's generalizability to new, candidate…

Machine Learning · Computer Science 2022-02-18 Leo Betthauser , Urszula Chajewska , Maurice Diesendruck , Rohith Pesala

Recent advances have shown that statistical tests for the rank of cross-covariance matrices play an important role in causal discovery. These rank tests include partial correlation tests as special cases and provide further graphical…

Machine Learning · Computer Science 2025-06-13 Xinshuai Dong , Ignavier Ng , Boyang Sun , Haoyue Dai , Guang-Yuan Hao , Shunxing Fan , Peter Spirtes , Yumou Qiu , Kun Zhang

We investigate the nonparametric, composite hypothesis testing problem for arbitrary unknown distributions in the asymptotic regime where both the sample size and the number of hypotheses grow exponentially large. Such asymptotic analysis…

Information Theory · Computer Science 2019-01-30 Qunwei Li , Tiexing Wang , Donald J. Bucci , Yingbin Liang , Biao Chen , Pramod K. Varshney

We consider semiparametric location-scatter models for which the $p$-variate observation is obtained as $X=\Lambda Z+\mu$, where $\mu$ is a $p$-vector, $\Lambda$ is a full-rank $p\times p$ matrix and the (unobserved) random $p$-vector $Z$…

Statistics Theory · Mathematics 2012-02-24 Pauliina Ilmonen , Davy Paindaveine

This paper studies the minimax detection of a small submatrix of elevated mean in a large matrix contaminated by additive Gaussian noise. To investigate the tradeoff between statistical performance and computational cost from a…

Statistics Theory · Mathematics 2015-06-04 Zongming Ma , Yihong Wu

Synchronized measurements of a large power grid enable an unprecedented opportunity to study the spatialtemporal correlations. Statistical analytics for those massive datasets start with high-dimensional data matrices. Uncertainty is…

Applications · Statistics 2018-02-13 Zenan Ling , Robert C. Qiu , Xing He , Lei Chu

We propose a new approach, the calibrated nonparametric scan statistic (CNSS), for more accurate detection of anomalous patterns in large-scale, real-world graphs. Scan statistics identify connected subgraphs that are interesting or…

Methodology · Statistics 2022-06-28 Chunpai Wang , Daniel B. Neill , Feng Chen

A common method for deriving non-parametric tests is to reformulate a parametric test in terms of sample ranks. Despite being distribution free (even in finite samples), the resulting tests often display remarkable asymptotic power…

Statistics Theory · Mathematics 2022-08-10 Dan D. Erdmann-Pham , Jonathan Terhorst , Yun S. Song

In this paper, we study the problem of detecting multiple hidden submatrices in a large Gaussian random matrix when the planted signal is inhomogeneous across entries. Under the null hypothesis, the observed matrix has independent and…

Statistics Theory · Mathematics 2026-03-13 Mor Oren-Loberman , Dvir Jerbi , Tamir Bendory , Wasim Huleihel

Invariance-based randomization tests -- such as permutation tests, rotation tests, or sign changes -- are an important and widely used class of statistical methods. They allow drawing inferences under weak assumptions on the data…

Statistics Theory · Mathematics 2022-05-31 Edgar Dobriban

Low-rank matrix completion concerns the problem of estimating unobserved entries in a matrix using a sparse set of observed entries. We consider the non-uniform setting where the observed entries are sampled with highly varying…

Machine Learning · Statistics 2024-03-04 Xumei Xi , Christina Lee Yu , Yudong Chen
‹ Prev 1 2 3 10 Next ›