中文
相关论文

相关论文: Two-sample Behrens--Fisher problems for high-dimen…

200 篇论文

Statistical data is often analyzed as a contingency table, sometimes with empty cells called zeros. Such sparse tables can be due to scarse observations classified in numerous categories, as for example in genetic association studies. Thus,…

统计理论 · 数学 2010-07-28 Audrey Finkler

Statistical data is often analyzed as a contingency table, sometimes with empty cells called zeros. Such sparse tables can be due to scarse observations classified in numerous categories, as for example in genetic association studies. Thus,…

统计理论 · 数学 2010-07-28 Audrey Finkler

The analysis of large-scale datasets, especially in biomedical contexts, frequently involves a principled screening of multiple hypotheses. The celebrated two-group model jointly models the distribution of the test statistics with mixtures…

统计方法学 · 统计学 2023-03-10 Francesco Denti , Stefano Peluso , Michele Guindani , Antonietta Mira

A family of maximum mean discrepancy (MMD) kernel two-sample tests is introduced. Members of the test family are called Block-tests or B-tests, since the test statistic is an average over MMDs computed on subsets of the samples. The choice…

机器学习 · 计算机科学 2014-02-11 Wojciech Zaremba , Arthur Gretton , Matthew Blaschko

For high-dimensional small sample size data, Hotelling's T2 test is not applicable for testing mean vectors due to the singularity problem in the sample covariance matrix. To overcome the problem, there are three main approaches in the…

统计方法学 · 统计学 2020-03-11 Zongliang Hu , Tiejun Tong , Marc G. Genton

In many real-world applications, it is common that a proportion of the data may be missing or only partially observed. We develop a novel two-sample testing method based on the Maximum Mean Discrepancy (MMD) which accounts for missing data…

统计方法学 · 统计学 2024-05-27 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

Due to the broad applications of elliptical models, there is a long line of research on goodness-of-fit tests for empirically validating them. However, the existing literature on this topic is generally confined to low-dimensional settings,…

统计理论 · 数学 2025-03-04 Siyao Wang , Miles E. Lopes

We consider the classification problem of a high-dimensional mixture of two Gaussians with general covariance matrices. Using the replica method from statistical physics, we investigate the asymptotic behavior of a general class of…

机器学习 · 统计学 2024-10-29 Hanwen Huang , Peng Zeng

Testing independence among a number of (ultra) high-dimensional random samples is a fundamental and challenging problem. By arranging $n$ identically distributed $p$-dimensional random vectors into a $p \times n$ data matrix, we investigate…

统计理论 · 数学 2017-03-28 Xi Chen , Weidong Liu

Two-sample hypothesis testing for network comparison presents many significant challenges, including: leveraging repeated network observations and known node registration, but without requiring them to operate; relaxing strong structural…

统计方法学 · 统计学 2024-02-05 Meijia Shao , Dong Xia , Yuan Zhang , Qiong Wu , Shuo Chen

Permutation tests date back nearly a century to Fisher's randomized experiments, and remain an immensely popular statistical tool, used for testing hypotheses of independence between variables and other common inferential questions. Much of…

统计方法学 · 统计学 2022-12-05 Aaditya Ramdas , Rina Foygel Barber , Emmanuel J. Candes , Ryan J. Tibshirani

High-dimensional logistic regression is widely used in analyzing data with binary outcomes. In this paper, global testing and large-scale multiple testing for the regression coefficients are considered in both single- and two-regression…

统计方法学 · 统计学 2020-11-23 Rong Ma , T. Tony Cai , Hongzhe Li

Testing for the equality of two high-dimensional distributions is a challenging problem, and this becomes even more challenging when the sample size is small. Over the last few decades, several graph-based two-sample tests have been…

统计方法学 · 统计学 2019-11-22 Soham Sarkar , Rahul Biswas , Anil K. Ghosh

We investigate the problem of testing the global null in the high-dimensional regression models when the feature dimension $p$ grows proportionally to the number of observations $n$. Despite a number of prior work studying this problem,…

统计方法学 · 统计学 2020-10-06 Yue Li , Ilmun Kim , Yuting Wei

Pearson's chi-squared test is widely used to test the goodness of fit between categorical data and a given discrete distribution function. When the number of sets of the categorical data, say $k$, is a fixed integer, Pearson's chi-squared…

统计方法学 · 统计学 2022-01-03 Shuhua Chang , Deli Li , Yongcheng Qi

We consider the problem of testing a null hypothesis defined by equality and inequality constraints on a statistical parameter. Testing such hypotheses can be challenging because the number of relevant constraints may be on the same order…

统计方法学 · 统计学 2024-02-19 Nils Sturma , Mathias Drton , Dennis Leung

The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification…

机器学习 · 统计学 2017-11-15 Hamed Masnadi-Shirazi

High-dimensional data clustering has become and remains a challenging task for modern statistics and machine learning, with a wide range of applications. We consider in this work the powerful discriminative latent mixture model, and we…

统计方法学 · 统计学 2020-12-09 Nicolas Jouvin , Charles Bouveyron , Pierre Latouche

Network data is a major object data type that has been widely collected or derived from common sources such as brain imaging. Such data contains numeric, topological, and geometrical information, and may be necessarily considered in certain…

统计方法学 · 统计学 2021-06-29 Han Feng , Xing Qiu , Hongyu Miao

We develop a general statistical framework for the analysis and inference of large tree-structured data, with a focus on developing asymptotic goodness-of-fit tests. We first propose a consistent statistical model for binary trees, from…

‹ 上一页 1 8 9 10 下一页 ›