Related papers: A Statistical Distance Derived From The Kolmogorov…
We investigate the statistics of the cosmic microwave background using the Kolmogorov-Smirnov test. We show that, when we correctly de-correlate the data, the partition function of the Kolmogorov stochasticity parameter is compatible with…
The assessment of segmentation quality plays a fundamental role in the development, optimization, and comparison of segmentation methods which are used in a wide range of applications. With few exceptions, quality assessment is performed…
We provide analytic formulas for the standard error and confidence intervals for the F measures, based on a property of asymptotic normality in the large sample limit. The formula can be applied for sample size planning in order to achieve…
Material microstructures are traditionally compared using sets of statistical measures that are incomplete, e.g., two visually distinct microstructures can have identical grain size distributions and phase fractions. While this is not a…
In this paper, the concept of the classical $f$-divergence (for a pair of measures) is extended to the mixed $f$-divergence (for multiple pairs of measures). The mixed $f$-divergence provides a way to measure the difference between multiple…
In various disordered systems or non-equilibrium dynamical models, the large deviations of some observables have been found to display different scalings for rare values bigger or smaller than the typical value. In the present paper, we…
In this paper, the concept of the classical $f$-divergence for a pair of measures is extended to the mixed $f$-divergence for multiple pairs of measures. The mixed $f$-divergence provides a way to measure the difference between multiple…
We prove the following three statements: 1) Let $(A, \bar A)$ be a partition of the spherical surface $S^n$ into two measurable sets. Let $st_A$ and $st_{\bar A}$ be their measure density functions of distance. Then $|st_A - st_{\bar A}|$…
Upon a consistent topological statistical theory the application of structural statistics requires a quantification of the proximity structure of model spaces. An important tool to study these structures are Pseudo-Riemannian metrices,…
This paper derives the rate of convergence and asymptotic distribution for a class of Kolmogorov-Smirnov style test statistics for conditional moment inequality models for parameters on the boundary of the identified set under general…
The use of descriptive statistics in pilot testing procedures requires objective, standard diagnostic tools that are feasible for small sample sizes. While current psychometric practices report item-level statistics, they often report these…
Poincar{\'e} inequalities are ubiquitous in probability and analysis and have various applications in statistics (concentration of measure, rate of convergence of Markov chains). The Poincar{\'e} constant, for which the inequality is tight,…
We discuss the role that the null hypothesis should play in the construction of a test statistic used to make a decision about that hypothesis. To construct the test statistic for a point null hypothesis about a binomial proportion, a…
A new class of distances appropriate for measuring similarity relations between sequences, say one type of similarity per distance, is studied. We propose a new ``normalized information distance'', based on the noncomputable notion of…
The F-measure or F-score is one of the most commonly used single number measures in Information Retrieval, Natural Language Processing and Machine Learning, but it is based on a mistake, and the flawed assumptions render it unsuitable for…
Friedman's chi-square test is a non-parametric statistical test for $r\geq2$ treatments across $n\ge1$ trials to assess the null hypothesis that there is no treatment effect. We use Stein's method with an exchangeable pair coupling to…
Most of the work on checking spherical symmetry assumptions on the distribution of the $p$-dimensional random vector $Y$ has its focus on statistical tests for the null hypothesis of exact spherical symmetry. In this paper, we take a…
Persistence diagrams (PDs) are used as signatures of point cloud data. Two clouds of points can be compared using the bottleneck distance d_B between their PDs. A potential drawback of this pipeline is that point clouds sampled from…
In machine learning, observation features are measured in a metric space to obtain their distance function for optimization. Given similar features that are statistically sufficient as a population, a statistical distance between two…
Transportation distance information is a powerful resource, but location records are often censored due to privacy concerns or regulatory mandates. We outline methods to approximate, sample from, and compare distributions of distances…