中文
相关论文

相关论文: B-tests: Low Variance Kernel Two-Sample Tests

200 篇论文

Maximum Mean Discrepancy (MMD) has been widely used in the areas of machine learning and statistics to quantify the distance between two distributions in the $p$-dimensional Euclidean space. The asymptotic property of the sample MMD has…

统计理论 · 数学 2023-08-29 Hanjia Gao , Xiaofeng Shao

Recent years have seen a surge in methods for two-sample testing, among which the Maximum Mean Discrepancy (MMD) test has emerged as an effective tool for handling complex and high-dimensional data. Despite its success and widespread…

机器学习 · 统计学 2026-05-21 Ikjun Choi , Ilmun Kim

We present a study of a kernel-based two-sample test statistic related to the Maximum Mean Discrepancy (MMD) in the manifold data setting, assuming that high-dimensional observations are close to a low-dimensional manifold. We characterize…

机器学习 · 统计学 2024-02-27 Xiuyuan Cheng , Yao Xie

Modern kernel-based two-sample tests have shown great success in distinguishing complex, high-dimensional distributions with appropriate learned kernels. Previous work has demonstrated that this kernel learning procedure succeeds, assuming…

机器学习 · 统计学 2022-01-06 Feng Liu , Wenkai Xu , Jie Lu , Danica J. Sutherland

The maximum mean discrepancy (MMD) is a kernel-based distance between probability distributions useful in many applications (Gretton et al. 2012), bearing a simple estimator with pleasing computational and statistical properties. Being able…

机器学习 · 统计学 2022-11-16 Danica J. Sutherland , Namrata Deka

Nonparametric two sample testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. We refer to the most common…

统计理论 · 数学 2015-08-05 Aaditya Ramdas , Sashank J. Reddi , Barnabas Poczos , Aarti Singh , Larry Wasserman

A new goodness-of-fit test for normality in high-dimension (and Reproducing Kernel Hilbert Space) is proposed. It shares common ideas with the Maximum Mean Discrepancy (MMD) it outperforms both in terms of computation time and applicability…

统计理论 · 数学 2014-04-14 Jérémie Kellner , Alain Celisse

Motivated by the increasing use of kernel-based metrics for high-dimensional and large-scale data, we study the asymptotic behavior of kernel two-sample tests when the dimension and sample sizes both diverge to infinity. We focus on the…

统计理论 · 数学 2024-10-31 Jian Yan , Xianyang Zhang

The Maximum Mean Discrepancy (MMD) is a cornerstone statistic for nonparametric two-sample testing, but its test power is dictated entirely by the chosen kernel. Because any fixed kernel inherently fails to distinguish certain…

机器学习 · 统计学 2026-05-11 Yijin Ni , Xiaoming Huo

In modern data analysis, nonparametric measures of discrepancies between random variables are particularly important. The subject is well-studied in the frequentist literature, while the development in the Bayesian setting is limited where…

统计方法学 · 统计学 2022-01-25 Qinyi Zhang , Veit Wild , Sarah Filippi , Seth Flaxman , Dino Sejdinovic

Nonparametric two-sample tests such as the Maximum Mean Discrepancy (MMD) are often used to detect differences between two distributions in machine learning applications. However, the majority of existing literature assumes that error-free…

机器学习 · 统计学 2023-08-08 Ron Nafshi , Maggie Makar

Kernel two-sample tests have been widely used for multivariate data to test equality of distributions. However, existing tests based on mapping distributions into a reproducing kernel Hilbert space mainly target specific alternatives and do…

统计方法学 · 统计学 2023-11-21 Hoseung Song , Hao Chen

In many real-world applications, it is common that a proportion of the data may be missing or only partially observed. We develop a novel two-sample testing method based on the Maximum Mean Discrepancy (MMD) which accounts for missing data…

统计方法学 · 统计学 2024-05-27 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

Maximum Mean Discrepancy (MMD) is a widely used concept in machine learning research which has gained popularity in recent years as a highly effective tool for comparing (finite-dimensional) distributions. Since it is designed as a…

机器学习 · 统计学 2025-06-03 Andrew Alden , Blanka Horvath , Zacharia Issa

We introduce a kernel-based goodness-of-fit test for censored data, where observations may be missing in random time intervals: a common occurrence in clinical trials and industrial life-testing. The test statistic is straightforward to…

统计方法学 · 统计学 2018-10-11 Tamara Fernández , Arthur Gretton

We characterize the asymptotic performance of nonparametric goodness of fit testing. The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, and a test is optimal if it achieves the maximum…

机器学习 · 统计学 2019-03-19 Shengyu Zhu , Biao Chen , Pengfei Yang , Zhitang Chen

Do two data samples come from different distributions? Recent studies of this fundamental problem focused on embedding probability distributions into sufficiently rich characteristic Reproducing Kernel Hilbert Spaces (RKHSs), to compare…

Two-sample tests have been extensively employed in various scientific fields and machine learning such as evaluation on the effectiveness of drugs and A/B testing on different marketing strategies to discriminate whether two sets of samples…

量子物理 · 物理学 2025-11-27 Yu Terada , Yugo Ogio , Ken Arai , Hiroyuki Tezuka , Yu Tanaka

We propose optimal Bayesian two-sample tests for testing equality of high-dimensional mean vectors and covariance matrices between two populations. In many applications including genomics and medical imaging, it is natural to assume that…

统计方法学 · 统计学 2021-12-07 Kyoungjae Lee , Kisung You , Lizhen Lin

We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, maximum mean…

统计方法学 · 统计学 2013-11-13 Dino Sejdinovic , Bharath Sriperumbudur , Arthur Gretton , Kenji Fukumizu