中文
相关论文

相关论文: A Tracy-Widom Empirical Estimator For Valid P-valu…

200 篇论文

We introduce principal differences analysis (PDA) for analyzing differences between high-dimensional distributions. The method operates by finding the projection that maximizes the Wasserstein divergence between the resulting univariate…

机器学习 · 统计学 2017-05-03 Jonas Mueller , Tommi Jaakkola

A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be…

机器学习 · 计算机科学 2019-09-12 Jonas Mueller , Alex Smola

This paper studies fundamental aspects of modelling data using multivariate Watson distributions. Although these distributions are natural for modelling axially symmetric data (i.e., unit vectors where $\pm \x$ are equivalent), for…

统计计算 · 统计学 2012-05-28 Suvrit Sra , Dmitrii Karp

Confidence intervals are a popular way to visualize and analyze data distributions. Unlike p-values, they can convey information both about statistical significance as well as effect size. However, very little work exists on applying…

应用统计 · 统计学 2017-01-23 Jussi Korpela , Emilia Oikarinen , Kai Puolamäki , Antti Ukkonen

Modern longitudinal studies collect multiple outcomes as the primary endpoints to understand the complex dynamics of the diseases. Oftentimes, especially in clinical trials, the joint variations among the multidimensional responses play a…

统计方法学 · 统计学 2024-01-17 Salil Koner , Sheng Luo

While the problem of testing multivariate normality has received considerable attention in the classical low-dimensional setting where the sample size $n$ is much larger than the feature dimension $d$ of the data, there is presently a…

统计方法学 · 统计学 2025-12-23 Xin Bing , Derek Latremouille

Repeated-measure designs allow comparisons within a group as well as between groups, and are commonly referred to as split-plot designs. While originating in agricultural experiments, they are now widely used in medical research,…

统计计算 · 统计学 2025-12-22 Paavo Sattler , Nils Hichert

Determining causal relationship between high dimensional observations are among the most important tasks in scientific discoveries. In this paper, we revisited the \emph{linear trace method}, a technique proposed…

机器学习 · 计算机科学 2023-03-15 Arun Jambulapati , Hilaf Hasson , Youngsuk Park , Yuyang Wang

In many settings, robust data analysis involves computational methods for uncertainty quantification and statistical inference. To design frequentist studies that leverage robust analysis methods, suitable sample sizes to achieve desired…

统计方法学 · 统计学 2025-12-19 Luke Hagar , Andrew J. Martin

Datasets containing both categorical and continuous variables are frequently encountered in many areas, and with the rapid development of modern measurement technologies, the dimensions of these variables can be very high. Despite the…

统计方法学 · 统计学 2024-01-03 Binyan Jiang , Chenlei Leng , Cheng Wang , Zhongqing Yang , Xinyang Yu

Detection of the number of signals corrupted by high-dimensional noise is a fundamental problem in signal processing and statistics. This paper focuses on a general setting where the high-dimensional noise has an unknown complicated…

统计理论 · 数学 2022-05-16 Xiucai Ding , Fan Yang

Multiway data analysis aims to uncover patterns in data structured as multi-indexed arrays, with multiway covariance playing a crucial role in many applications. However, the high dimensionality of multiway covariance presents significant…

统计理论 · 数学 2026-03-19 Dogyoon Song , Alfred O. Hero

We study the sample covariance matrix for real-valued data with general population covariance, as well as MANOVA-type covariance estimators in variance components models under null hypotheses of global sphericity. In the limit as matrix…

概率论 · 数学 2020-06-11 Zhou Fan , Iain M. Johnstone

This article inspects whether a multivariate distribution is different from a specified distribution or not, and it also tests the equality of two multivariate distributions. In the course of this study, a graphical tool-kit using…

统计方法学 · 统计学 2024-08-19 Pratim Guha Niyogi , Subhra Sankar Dhar

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

统计方法学 · 统计学 2024-09-05 F. Richard Guo , Rajen D. Shah

In this paper, we develop new statistical theory for probabilistic principal component analysis models in high dimensions. The focus is the estimation of the noise variance, which is an important and unresolved issue when the number of…

统计理论 · 数学 2014-06-23 Damien Passemier , Zhaoyuan Li , Jian-Feng Yao

A ubiquitous feature of data of our era is their extra-large sizes and dimensions. Analyzing such high-dimensional data poses significant challenges, since the feature dimension is often much larger than the sample size. This thesis…

统计理论 · 数学 2025-09-11 Kai Yang

We consider the hypothesis testing problem of detecting a shift between the means of two multivariate normal distributions in the high-dimensional setting, allowing for the data dimension p to exceed the sample size n. Specifically, we…

统计理论 · 数学 2015-09-15 Miles E. Lopes , Laurent J. Jacob , Martin J. Wainwright

We studied the universality of Wishart ensembles whose covariance matrix has 2 distinct eigenvalues. We studied the asymptotic limit when the number of both eigenvalues goes to infinity and obtained universality results. In this case, the…

概率论 · 数学 2008-09-26 M. Y. Mo

Evaluating the joint significance of covariates is of fundamental importance in a wide range of applications. To this end, p-values are frequently employed and produced by algorithms that are powered by classical large-sample asymptotic…

统计方法学 · 统计学 2017-05-11 Yingying Fan , Emre Demirkaya , Jinchi Lv