English
Related papers

Related papers: A Tracy-Widom Empirical Estimator For Valid P-valu…

200 papers

High-dimensional vector autoregressive (VAR) models provide a flexible framework for characterizing dynamic dependence in multivariate spatio-temporal systems, but their unrestricted estimation becomes infeasible when multiple variables are…

Methodology · Statistics 2026-05-04 Peiliang Bai

Distributional approximations of (bi--) linear functions of sample variance-covariance matrices play a critical role to analyze vector time series, as they are needed for various purposes, especially to draw inference on the dependence…

Probability · Mathematics 2018-03-20 Ansgar Steland , Rainer von Sachs

In this paper, we propose a data-adaptive empirical likelihood-based approach for treatment effect estimation and inference, which overcomes the obstacle of the traditional empirical likelihood-based approaches in the high-dimensional…

Methodology · Statistics 2020-12-15 Wei Liang , Ying Yan

A novel method for common and individual feature analysis from exceedingly large-scale data is proposed, in order to ensure the tractability of both the computation and storage and thus mitigate the curse of dimensionality, a major…

Signal Processing · Electrical Eng. & Systems 2017-11-03 Ilia Kisil , Giuseppe G. Calvi , Danilo P. Mandic

Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically…

Machine Learning · Computer Science 2025-02-25 Ya-Wei Eileen Lin , Ronald R. Coifman , Gal Mishne , Ronen Talmon

Applications of high-dimensional regression often involve multiple sources or types of covariates. We propose methodology for this setting, emphasizing the "wide data" regime with large total dimensionality p and sample size n<<p. We focus…

This paper develops new tools to quantify uncertainty in optimal decision making and to gain insight into which variables one should collect information about given the potential cost of measuring a large number of variables. We investigate…

Methodology · Statistics 2021-05-11 Yunan Wu , Lan Wang , Haoda Fu

Let $(X,Y)$ be a random variable consisting of an observed feature vector $X\in \mathcal{X}$ and an unobserved class label $Y\in \{1,2,...,L\}$ with unknown joint distribution. In addition, let $\mathcal{D}$ be a training data set…

Statistics Theory · Mathematics 2008-06-26 Lutz Duembgen , Bernd-Wolfgang Igl , Axel Munk

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional statistical methods often fail to produce sensible results due…

Methodology · Statistics 2016-01-19 Yuanpei Cao , Wei Lin , Hongzhe Li

Matrix-variate distributions can intuitively model the dependence structure of matrix-valued observations that arise in applications with multivariate time series, spatio-temporal or repeated measures. This paper develops an…

Methodology · Statistics 2019-12-24 Geoffrey Z. Thompson , Ranjan Maitra , William Q. Meeker , Ashraf Bastawros

In this paper, we consider the problem of deriving new eigenvalue distributions of real-valued Wishart matrices that arises in many scientific and engineering applications. The distributions are derived using the tools from the theory of…

Information Theory · Computer Science 2015-07-29 Oliver James , Heung-No Lee

Two-sample hypothesis testing is a fundamental problem with various applications, which faces new challenges in the high-dimensional context. To mitigate the issue of the curse of dimensionality, high-dimensional data are typically assumed…

Methodology · Statistics 2026-04-06 Jiaqi Gu , Ruoxu Tan , Guosheng Yin

The classic Hettmansperger-Randles Estimator has found extensive use in robust statistical inference. However, it cannot be directly applied to high-dimensional data. In this paper, we propose a high-dimensional Hettmansperger-Randles…

Methodology · Statistics 2025-05-06 Guowei Yan , Long Feng , Xiaoxu Zhang

Let $\bY =\bR+\bX$ be an $M\times N$ matrix, where $\bR$ is a rectangular diagonal matrix and $\bX$ consists of $i.i.d.$ entries. This is a signal-plus-noise type model. Its signal matrix could be full rank, which is rarely studied in…

Statistics Theory · Mathematics 2020-09-28 Zhixiang Zhang , Guangming Pan

Our data are random fields of multivariate Gaussian observations, and we fit a multivariate linear model with common design matrix at each point. We are interested in detecting those points where some of the coefficients are nonzero using…

Statistics Theory · Mathematics 2009-09-29 J. E. Taylor , K. J. Worsley

We investigate the random eigenvalues coming from the beta-Laguerre ensemble with parameter p, which is a generalization of the real, complex and quaternion Wishart matrices of parameter (n,p). In the case that the sample size n is much…

Probability · Mathematics 2013-09-17 Tiefeng Jiang , Danning Li

In this work, we consider causal inference in various high-dimensional treatment settings, including for single multi-valued treatments and vector treatments with binary or continuous components, when the number of treatments can be…

Statistics Theory · Mathematics 2026-02-26 Patrick Kramer , Edward H. Kennedy , Isaac M. Opper

This paper extends the work of El Karoui [Ann. Probab. 35 (2007) 663--714] which finds the Tracy--Widom limit for the largest eigenvalue of a nonsingular $p$-dimensional complex Wishart matrix $W_{\mathbb{C}}(\Omega_p,n)$ to the case of…

Probability · Mathematics 2008-12-18 Alexei Onatski

Poor performance of quantitative analysis in histopathological Whole Slide Images (WSI) has been a significant obstacle in clinical practice. Annotating large-scale WSIs manually is a demanding and time-consuming task, unlikely to yield the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Sarah Cechnicka , James Ball , Hadrien Reynaud , Callum Arthurs , Candice Roufosse , Bernhard Kainz

High-dimensional data arise routinely in modern statistics, econometrics, finance, genomics, and machine learning. While a large body of existing methodology is developed under Gaussian or light-tailed assumptions, many real data sets…

Methodology · Statistics 2026-04-16 Long Feng