English
Related papers

Related papers: Distances between Data Sets Based on Summary Stati…

200 papers

A new class of distances appropriate for measuring similarity relations between sequences, say one type of similarity per distance, is studied. We propose a new ``normalized information distance'', based on the noncomputable notion of…

Computational Complexity · Computer Science 2011-11-09 Ming Li , Xin Chen , Xin Li , Bin Ma , Paul Vitanyi

The average distance from a node to all other nodes in a graph, or from a query point in a metric space to a set of points, is a fundamental quantity in data analysis. The inverse of the average distance, known as the (classic) closeness…

Social and Information Networks · Computer Science 2015-06-29 Shiri Chechik , Edith Cohen , Haim Kaplan

This paper is about similarity between objects that can be represented as points in metric measure spaces. A metric measure space is a metric space that is also equipped with a measure. For example, a network with distances between its…

Discrete Mathematics · Computer Science 2020-11-03 Evgeny Dantsin , Alexander Wolpert

Statistical distances quantifies the difference between two statistical constructs. In this article, we describe reference values for a distance between samples derived from the Kolmogorov-Smirnov statistic $D_{F,F'}$. Each measure of the…

Data Analysis, Statistics and Probability · Physics 2017-11-03 Renato Fabbri , Fernando Gularte De León

For better learning, large datasets are often split into small batches and fed sequentially to the predictive model. In this paper, we study such batch decompositions from a probabilistic perspective. We assume that data points (possibly…

Machine Learning · Computer Science 2025-04-10 Ghurumuruhan Ganesan

We provide an implementation to compute the flat metric in any dimension. The flat metric, also called dual bounded Lipschitz distance, generalizes the well-known Wasserstein distance $W_1$ to the case that the distributions are of unequal…

Machine Learning · Computer Science 2025-06-17 Henri Schmidt , Christian Düll

This work builds a unified framework for the study of quadratic form distance measures as they are used in assessing the goodness of fit of models. Many important procedures have this structure, but the theory for these methods is dispersed…

Statistics Theory · Mathematics 2008-12-18 Bruce G. Lindsay , Marianthi Markatou , Surajit Ray , Ke Yang , Shu-Chuan Chen

Many learning algorithms such as kernel machines, nearest neighbors, clustering, or anomaly detection, are based on the concept of 'distance' or 'similarity'. Before similarities are used for training an actual machine learning model, we…

We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize…

This paper defines an alternative notion, described as data-based, of geometric quantiles on Hadamard spaces, in contrast to the existing methodology, described as parameter-based. In addition to having the same desirable properties as…

Methodology · Statistics 2025-06-17 Ha-Young Shin , Hee-Seok Oh

The Levenshtein distance is an important tool for the comparison of symbolic sequences, with many appearances in genome research, linguistics and other areas. For efficient applications, an approximation by a distance of smaller…

Quantitative Methods · Quantitative Biology 2007-05-23 Michael Baake , Uwe Grimm , Robert Giegerich

Measuring the distance between concepts is an important field of study of Natural Language Processing, as it can be used to improve tasks related to the interpretation of those same concepts. WordNet, which includes a wide variety of…

Computation and Language · Computer Science 2018-04-30 Raquel Pérez-Arnal , Armand Vilalta , Dario Garcia-Gasulla , Ulises Cortés , Eduard Ayguadé , Jesus Labarta

The emergence of time-series foundation model research elevates the growing need to measure the (dis)similarity of time-series datasets. A time-series dataset similarity measure aids research in multiple ways, including model selection,…

Machine Learning · Computer Science 2025-07-31 Hongjie Chen , Akshay Mehra , Josh Kimball , Ryan A. Rossi

We define two minimum distance estimators for dependent data by minimizing some approximated Maximum Mean Discrepancy distances between the true empirical distribution of observations and their assumed (parametric) model distribution. When…

Methodology · Statistics 2026-01-19 Pierre Alquier , Jean-David Fermanian , Benjamin Poignard

While Kolmogorov complexity is the accepted absolute measure of information content in an individual finite object, a similarly absolute notion is needed for the information distance between two individual objects, for example, two…

Information Theory · Computer Science 2010-06-18 Charles H. Bennett , Peter Gacs , Ming Li , Paul M. B. Vitanyi , Wojciech H. Zurek

Can the spatial distance between two identical particles be explained in terms of the extent that one can be distinguished from the other? Is the geometry of space a macroscopic manifestation of an underlying microscopic statistical…

General Relativity and Quantum Cosmology · Physics 2015-06-25 Ariel Caticha

The similarity between objects is significant in a broad range of areas. While similarity can be measured using off-the-shelf distance functions, they may fail to capture the inherent meaning of similarity, which tends to depend on the…

Quantum Physics · Physics 2022-01-10 Santosh Kumar Radha , Casey Jao

Comparing datasets is a fundamental task in machine learning, essential for various learning paradigms-from evaluating train and test datasets for model generalization to using dataset similarity for detecting data drift. While traditional…

Machine Learning · Computer Science 2025-06-18 Paula Rodriguez-Diaz , Lingkai Kong , Kai Wang , David Alvarez-Melis , Milind Tambe

Classical machine learning approaches are sensitive to non-stationarity. Transfer learning can address non-stationarity by sharing knowledge from one system to another, however, in areas like machine prognostics and defense, data is…

Machine Learning · Computer Science 2022-09-07 Tyler Cody , Stephen Adams , Peter A. Beling

Distance function is a main metrics of measuring the affinity between two data points in machine learning. Extant distance functions often provide unreachable distance values in real applications. This can lead to incorrect measure of the…

Machine Learning · Computer Science 2022-07-14 Shichao Zhang , Jiaye Li , Yangding Li