Related papers: Testing to distinguish measures on metric spaces
We show that the Gamma distribution is not an adequate fit for the probability density function of drop diameters using the Kolmogorov-Smirnov goodness of fit test. We propose a different parametrization of drop size distributions, which…
The two-sample Kolmogorov-Smirnov test is a widely used statistical test for detecting whether two samples are likely to come from the same distribution. Implementations typically recur on an article of Hodges from 1957. The advances in…
We propose an application of the Kolmogorov-Smirnov test for rapidity distributions of individual events in ultrarelativistic heavy ion collisions. The test is particularly suitable to recognise non-statistical differences between the…
The field of property testing of probability distributions, or distribution testing, aims to provide fast and (most likely) correct answers to questions pertaining to specific aspects of very large datasets. In this work, we consider a…
Given samples from two distributions over an $n$-element set, we wish to test whether these distributions are statistically close. We present an algorithm which uses sublinear in $n$, specifically, $O(n^{2/3}\epsilon^{-8/3}\log n)$,…
Given a metric space with a Borel probability measure, for each integer $N$ we obtain a probability distribution on $N\times N$ distance matrices by considering the distances between pairs of points in a sample consisting of $N$ points…
We study here the error of numerical integration on metric measure spaces adapted to a decomposition of the space into disjoint subsets. We consider both the error for a single given function, and the worst case error for all functions in a…
The Gromov-Hausdorff distance measures the similarity between two metric spaces by isometrically embedding them into an ambient metric space. We introduce an analogue of this distance for metric spaces endowed with directed structures. The…
Kernel embeddings of distributions and the Maximum Mean Discrepancy (MMD), the resulting distance between distributions, are useful tools for fully nonparametric two-sample testing and learning on distributions. However, it is rarely that…
This paper investigates the estimation of the self-similarity parameter in fractional processes. We re-examine the Kolmogorov-Smirnov (KS) test as a distribution-based method for assessing self-similarity, emphasizing its robustness and…
A new class of distances appropriate for measuring similarity relations between sequences, say one type of similarity per distance, is studied. We propose a new ``normalized information distance'', based on the noncomputable notion of…
We consider a system of weak* closed sets of finite-dimensional distributions. We show that a corresponding system of random variables can be defined on a probability space with a probability measure determined up to some set of measures,…
We derive a new discrepancy statistic for measuring differences between two probability distributions based on combining Stein's identity with the reproducing kernel Hilbert space theory. We apply our result to test how well a probabilistic…
In statistics permutations typically arise in the context of rank plots for two-dimensional data. Such plots can also be interpreted as discrete copulas. In discrete mathematics, typically in the context of the description of large…
In this paper, the concept of the classical $f$-divergence (for a pair of measures) is extended to the mixed $f$-divergence (for multiple pairs of measures). The mixed $f$-divergence provides a way to measure the difference between multiple…
We propose a class of nonparametric two-sample tests with a cost linear in the sample size. Two tests are given, both based on an ensemble of distances between analytic functions representing each of the distributions. The first test uses…
We study the question of identity testing for structured distributions. More precisely, given samples from a {\em structured} distribution $q$ over $[n]$ and an explicit distribution $p$ over $[n]$, we wish to distinguish whether $q=p$…
Maximum Mean Discrepancy (MMD) is a widely used concept in machine learning research which has gained popularity in recent years as a highly effective tool for comparing (finite-dimensional) distributions. Since it is designed as a…
We propose a simple way of testing whether a given set of observations can come from a given theoretical cumulative distribution. In the test more weight is attached to the tails of the distribution than in the usual Kolmogorov or Smirnov…
Consider a set of multivariate distributions, $F_1,\dots,F_M$, aiming to explain the same phenomenon. For instance, each $F_m$ may correspond to a different candidate background model for calibration data, or to one of many possible signal…