中文
相关论文

相关论文: How many moments does MMD compare?

200 篇论文

We introduce a kernel-based two-sample test for comparing probability distributions up to group actions. Our construction yields invariant kernels for locally compact $\sigma$-compact groups and extends classical Haar-based approaches…

统计理论 · 数学 2026-03-18 Madison Giacofci , Anouar Meynaoui , Alex Podgorny

Nonparametric two-sample tests such as the Maximum Mean Discrepancy (MMD) are often used to detect differences between two distributions in machine learning applications. However, the majority of existing literature assumes that error-free…

机器学习 · 统计学 2023-08-08 Ron Nafshi , Maggie Makar

By facilitating the generation of samples from arbitrary probability distributions, Markov Chain Monte Carlo (MCMC) is, arguably, \emph{the} tool for the evaluation of Bayesian inference problems that yield non-standard posterior…

统计计算 · 统计学 2021-05-27 Peter L Green , Robert E Moore , Ryan J Jackson , Jinglai Li , Simon Maskell

Likelihood-free inference methods typically make use of a distance between simulated and real data. A common example is the maximum mean discrepancy (MMD), which has previously been used for approximate Bayesian computation, minimum…

统计方法学 · 统计学 2023-05-11 Ayush Bharti , Masha Naslidnyk , Oscar Key , Samuel Kaski , François-Xavier Briol

The maximum mean discrepancy (MMD) is a recently proposed test statistic for two-sample test. Its quadratic time complexity, however, greatly hampers its availability to large-scale applications. To accelerate the MMD calculation, in this…

人工智能 · 计算机科学 2015-06-19 Ji Zhao , Deyu Meng

We reprove the well known fact that the energy distance defines a metric on the space of Borel probability measures on a Hilbert space with finite first moment by a new approach, by analyzing the behavior of the Gaussian kernel on Hilbert…

泛函分析 · 数学 2021-02-02 Jean Carlo Guella

The ability to identify useful features or representations of the input data based on training data that achieves low prediction error on test data across multiple prediction tasks is considered the key to multitask learning success. In…

机器学习 · 统计学 2025-02-12 Soumya Mukherjee , Bharath K. Sriperumbudur

Divergences are quantities that measure discrepancy between two probability distributions and play an important role in various fields such as statistics and machine learning. Divergences are non-negative and are equal to zero if and only…

统计理论 · 数学 2019-10-22 Tomohiro Nishiyama

Do two data samples come from different distributions? Recent studies of this fundamental problem focused on embedding probability distributions into sufficiently rich characteristic Reproducing Kernel Hilbert Spaces (RKHSs), to compare…

The regularity of integration kernels forces decay rates of singular values of associated integral operators. This is well-known for symmetric operators with kernels defined on $(a,b) \times (a,b)$, where $(a,b)$ is an interval. Over time,…

泛函分析 · 数学 2023-11-07 Darko Volkov

In the theory of singular integral operators significant effort is often required to rigorously define such an operator. This is due to the fact that the kernels of such operators are not locally integrable on the diagonal, so the integral…

经典分析与常微分方程 · 数学 2014-03-31 Constanze Liaw , Sergei Treil

We propose a novel kernel-based two-sample test that leverages the spectral decomposition of the maximum mean discrepancy (MMD) statistic to identify and utilize well-estimated directional components in reproducing kernel Hilbert space…

统计方法学 · 统计学 2025-08-21 Rui Cui , Yuhao Li , Xiaojun Song

Machine learning (ML) models, such as SVM, for tasks like classification and clustering of sequences, require a definition of distance/similarity between pairs of sequences. Several methods have been proposed to compute the similarity…

Mean embeddings provide an extremely flexible and powerful tool in machine learning and statistics to represent probability distributions and define a semi-metric (MMD, maximum mean discrepancy; also called N-distance or energy distance),…

机器学习 · 统计学 2019-05-17 Matthieu Lerasle , Zoltan Szabo , Timothee Mathieu , Guillaume Lecue

Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…

机器学习 · 统计学 2026-04-21 Antoine Chatalic , Marco Letizia , Nicolas Schreuder , Lorenzo Rosasco

Mixed-integer linear programming (MILP) is a powerful tool for addressing a wide range of real-world problems, but it lacks a clear structure for comparing instances. A reliable similarity metric could establish meaningful relationships…

机器学习 · 计算机科学 2025-07-16 Gwen Maudet , Grégoire Danoy

Several researchers have proposed minimisation of maximum mean discrepancy (MMD) as a method to quantise probability measures, i.e., to approximate a target distribution by a representative point set. We consider sequential algorithms that…

机器学习 · 统计学 2021-02-15 Onur Teymur , Jackson Gorham , Marina Riabiz , Chris. J. Oates

The Earth Mover's Distance (EMD) is a state-of-the art metric for comparing discrete probability distributions, but its high distinguishability comes at a high cost in computational complexity. Even though linear-complexity approximation…

机器学习 · 计算机科学 2019-05-29 Kubilay Atasu , Thomas Mittelholzer

Existing measures and representations for trajectories have two longstanding fundamental shortcomings, i.e., they are computationally expensive and they can not guarantee the `uniqueness' property of a distance function: dist(X,Y) = 0 if…

机器学习 · 计算机科学 2023-01-03 Yufan Wang , Kai Ming Ting , Yuanyi Shang

We define and study pseudo-differential operators on a class of fractals that include the post-critically finite self-similar sets and Sierpinski carpets. Using the sub-Gaussian estimates of the heat operator we prove that our operators…

泛函分析 · 数学 2012-07-31 Marius Ionescu , Luke G. Rogers , Robert S. Strichartz