English
Related papers

Related papers: A Tracy-Widom Empirical Estimator For Valid P-valu…

200 papers

Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait.…

Machine Learning · Statistics 2012-05-31 Chamont Wang , Jana Gevertz , Chaur-Chin Chen , Leonardo Auslender

We consider settings where the observations are drawn from a zero-mean multivariate (real or complex) normal distribution with the population covariance matrix having eigenvalues of arbitrary multiplicity. We assume that the eigenvectors of…

Statistics Theory · Mathematics 2009-01-22 N. Raj Rao , James A. Mingo , Roland Speicher , Alan Edelman

Modern data often take the form of a multiway array. However, most classification methods are designed for vectors, i.e., 1-way arrays. Distance weighted discrimination (DWD) is a popular high-dimensional classification method that has been…

Methodology · Statistics 2021-10-12 Bin Guo , Lynn E. Eberly , Pierre-Gilles Henry , Christophe Lenglet , Eric F. Lock

Genome-wide association studies (GWAS) have led to the discovery of numerous single nucleotide polymorphisms (SNPs) associated with various phenotypes and complex diseases. However, the identified genetic variants do not fully explain the…

Methodology · Statistics 2025-07-09 Dayeon Jung , Yewon Kim , Junyong Park

Recent likelihood theory produces $p$-values that have remarkable accuracy and wide applicability. The calculations use familiar tools such as maximum likelihood values (MLEs), observed information and parameter rescaling. The usual…

Methodology · Statistics 2008-02-08 M. Bédard , D. A. S. Fraser , A. Wong

This article considers a novel and widely applicable approach to modeling high-dimensional dependent data when a large number of explanatory variables are available and the signal-to-noise ratio is low. We postulate that a $p$-dimensional…

Methodology · Statistics 2024-12-09 Zhaoxing Gao , Ruey S. Tsay

Handling multiplicity without losing much power has been a persistent challenge in various fields that often face the necessity of managing numerous statistical tests simultaneously. Recently, $p$-value combination methods based on…

Statistics Theory · Mathematics 2024-02-06 Yeonwoo Rho

Big Data involves both a large number of events but also many variables. This paper will concentrate on the challenge presented by the large number of variables in a Big Dataset. It will start with a brief review of exploratory data…

Applications · Statistics 2019-07-24 S. J. Watts , L. Crow

High-dimensional data can be useful for causal inference by providing many confounders that may bolster the plausibility of the ignorability assumption. Propensity score methods are powerful tools for causal inference, are popular in health…

Methodology · Statistics 2017-10-10 Jacob Spertus , Sharon-Lise Normand

Network meta-analysis of diagnostic test accuracy (NMA-DTA) is a relatively new field, involving combining evidence across studies to evaluate and compare the accuracy of different tests for a given condition. However, the methods proposed…

Methodology · Statistics 2026-04-23 Efthymia Derezea , Gabriel Rogers , Nicky J Welton , Hayley E Jones

In this paper, we first briefly review some recent results on the distribution of the maximal eigenvalue of a $(N\times N)$ random matrix drawn from Gaussian ensembles. Next we focus on the Gaussian Unitary Ensemble (GUE) and by suitably…

Statistical Mechanics · Physics 2011-05-30 Celine Nadal , Satya N. Majumdar

In many important statistical analyses, the number of covariates $p$ often exceeds the data size $n$, a regime commonly referred to as high-dimensional. While considerable progress has been made in high-dimensional regression under the…

Methodology · Statistics 2026-05-29 Herman Tesso , Georges Nguefack-Tsague

Testing independence among a number of (ultra) high-dimensional random samples is a fundamental and challenging problem. By arranging $n$ identically distributed $p$-dimensional random vectors into a $p \times n$ data matrix, we investigate…

Statistics Theory · Mathematics 2017-03-28 Xi Chen , Weidong Liu

We study simultaneous inference for multiple matrix-variate Gaussian graphical models in high-dimensional settings. Such models arise when spatiotemporal data are collected across multiple sample groups or experimental sessions, where each…

Methodology · Statistics 2026-01-21 Zongge Liu , Heejong Bong , Zhao Ren , Matthew A. Smith , Robert E. Kass

The correlated Wishart model provides the standard benchmark when analyzing time series of any kind. Unfortunately, the real case, which is the most relevant one in applications, poses serious challenges for analytical calculations. Often…

Mathematical Physics · Physics 2018-08-08 Tim Wirtz , Mario Kieburg , Thomas Guhr

How does one find dimensions in multivariate data that are reliably expressed across repetitions? For example, in a brain imaging study one may want to identify combinations of neural signals that are reliably expressed across multiple…

Machine Learning · Statistics 2022-12-05 Lucas C. Parra , Stefan Haufe , Jacek P. Dmochowski

Several recently developed methods have the potential to harness machine learning in the pursuit of target quantities inspired by causal inference, including inverse weighting, doubly robust estimating equations and substitution estimators…

The paper "An efficient sampling scheme for the eigenvalues of dual Wishart matrices", by I.~Santamar\'ia and V.~Elvira, [\emph{IEEE Signal Processing Letters}, vol.~28, pp.~2177--2181, 2021] \cite{SE21}, poses the question of efficient…

Statistics Theory · Mathematics 2024-01-24 Peter J. Forrester

Multi-dimensional data frequently occur in many different fields, including risk management, insurance, biology, environmental sciences, and many more. In analyzing multivariate data, it is imperative that the underlying modelling…

Methodology · Statistics 2025-06-23 Orla A. Murphy , Juliana Schulz

In many applications, linear models fit the data poorly. This article studies an appealing alternative, the generalized regression model. This model only assumes that there exists an unknown monotonically increasing link function connecting…

Methodology · Statistics 2017-07-24 Fang Han , Hongkai Ji , Zhicheng Ji , Honglang Wang