English
Related papers

Related papers: A Martingale Kernel Two-Sample Test

200 papers

In many practical transfer learning scenarios, the feature distribution is different across the source and target domains (i.e. non-i.i.d.). Maximum mean discrepancy (MMD), as a domain discrepancy metric, has achieved promising performance…

Computer Vision and Pattern Recognition · Computer Science 2019-03-26 Lei Zhang , Shanshan Wang , Guang-Bin Huang , Wangmeng Zuo , Jian Yang , David Zhang

In this paper, we propose a test for the equality of multiple distributions based on kernel mean embeddings. Our framework provides a flexible way to handle multivariate or even high-dimensional data by virtue of kernel methods and allows…

Statistics Theory · Mathematics 2020-06-08 Ilmun Kim

We propose a sequential test for detecting arbitrary distribution shifts that allows conformal test martingales (CTMs) to work under a fixed, reference-conditional setting. Existing CTM detectors construct test martingales by continually…

Machine Learning · Computer Science 2026-02-17 Shalev Shaer , Yarin Bar , Drew Prinster , Yaniv Romano

The kernel mean embedding of probability distributions is commonly used in machine learning as an injective mapping from distributions to functions in an infinite dimensional Hilbert space. It allows us, for example, to define a distance…

Quantum Physics · Physics 2019-12-24 Jonas M. Kübler , Krikamol Muandet , Bernhard Schölkopf

This article considers the popular MCMC method of unadjusted Langevin Monte Carlo (LMC) and provides a non-asymptotic analysis of its sampling error in 2-Wasserstein distance. The proof is based on a refinement of mean-square analysis in Li…

Machine Learning · Computer Science 2022-02-22 Ruilin Li , Hongyuan Zha , Molei Tao

One of the most well-known and simplest models for diversity maximization is the Max-Min Diversification (MMD) model, which has been extensively studied in the data mining and database literature. In this paper, we initiate the study of the…

Data Structures and Algorithms · Computer Science 2025-02-05 Iiro Kumpulainen , Florian Adriaens , Nikolaj Tatti

This paper is motivated by addressing open questions in distributionally robust chance-constrained programs (DRCCP) using the popular Wasserstein ambiguity sets. Specifically, the computational techniques for those programs typically place…

Optimization and Control · Mathematics 2022-04-26 Yassine Nemmour , Heiner Kremer , Bernhard Schölkopf , Jia-Jie Zhu

The Minimum Covariance Determinant (MCD) method is a highly robust estimator of multivariate location and scatter, for which a fast algorithm is available. Since estimating the covariance matrix is the cornerstone of many multivariate…

Methodology · Statistics 2021-01-13 Mia Hubert , Michiel Debruyne , Peter J. Rousseeuw

The statistics and machine learning communities have recently seen a growing interest in classification-based approaches to two-sample testing. The outcome of a classification-based two-sample test remains a rejection decision, which is not…

Statistics Theory · Mathematics 2022-11-15 Loris Michel , Jeffrey Näf , Nicolai Meinshausen

Many works in statistics aim at designing a universal estimation procedure, that is, an estimator that would converge to the best approximation of the (unknown) data generating distribution in a model, without any assumption on this…

Statistics Theory · Mathematics 2025-02-14 Badr-Eddine Chérief-Abdellatif , Pierre Alquier

We introduce a kernel-based goodness-of-fit test for censored data, where observations may be missing in random time intervals: a common occurrence in clinical trials and industrial life-testing. The test statistic is straightforward to…

Methodology · Statistics 2018-10-11 Tamara Fernández , Arthur Gretton

We study two-sample variable selection: identifying variables that discriminate between the distributions of two sets of data vectors. Such variables help scientists understand the mechanisms behind dataset discrepancies. Although…

Machine Learning · Statistics 2025-11-06 Kensuke Mitsuzawa , Motonobu Kanagawa , Stefano Bortoli , Margherita Grossi , Paolo Papotti

A new robust pairwise statistic, the pairwise median scaled difference (MSD), is proposed for the detection of anomalous location/uncertainty pairs in heteroscedastic interlaboratory study data with associated uncertainties. The…

Applications · Statistics 2018-10-05 Stephen L. R. Ellison

Reproducing Kernel Hilbert Space (RKHS) embedding of probability distributions has proved to be an effective approach, via MMD (maximum mean discrepancy), for nonparametric hypothesis testing problems involving distributions defined over…

Statistics Theory · Mathematics 2025-10-17 Soumya Mukherjee , Bharath K. Sriperumbudur

Comparing $K$-sample distributions is a fundamental problem in data science that arises in a wide variety of fields and applications. In this article, we introduce a maximum-of-differences approach to make such comparisons. Specifically, we…

Methodology · Statistics 2026-04-13 Wei Lan , Long Feng , Runze Li , Chih-Ling Tsai

The learning of domain-invariant representations in the context of domain adaptation with neural networks is considered. We propose a new regularization method that minimizes the discrepancy between domain-specific latent feature…

Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data,…

Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample…

Machine Learning · Statistics 2024-06-27 Cencheng Shen , Joshua T. Vogelstein

Stein showed that the multivariate sample mean is outperformed by "shrinking" to a constant target vector. Ledoit and Wolf extended this approach to the sample covariance matrix and proposed a multiple of the identity as shrinkage target.…

Methodology · Statistics 2014-12-08 Daniel Bartz , Johannes Höhne , Klaus-Robert Müller

We propose a novel method to learn intractable distributions from their samples. The main idea is to use a parametric distribution model, such as a Gaussian Mixture Model (GMM), to approximate intractable distributions by minimizing the…

Machine Learning · Computer Science 2023-08-15 Chenqiu Zhao , Guanfang Dong , Anup Basu