English
Related papers

Related papers: Distances for Comparing Multisets and Sequences

200 papers

This article provides an overview on the statistical modeling of complex data as increasingly encountered in modern data analysis. It is argued that such data can often be described as elements of a metric space that satisfies certain…

Methodology · Statistics 2024-02-28 Paromita Dubey , Yaqing Chen , Hans-Georg Müller

We modeled the dynamics of a soccer match based on a network representation where players are nodes discretely clustered into homogeneous groups. Players were grouped by physical proximity, supported by the intuitive notion that competing…

Social and Information Networks · Computer Science 2021-08-09 Luis Ramada Pereira , Rui J. Lopes , Jorge Louçã , Duarte Araújo , João Ramos

The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two…

Machine Learning · Computer Science 2025-09-24 Varun Babbar , Zhicheng Guo , Cynthia Rudin

Consider an unlimited homogeneous medium disturbed by points generated via Poisson process. The neighborhood of a point plays an important role in spatial statistics problems. Here, we obtain analytically the distance statistics to $k$th…

Statistical Mechanics · Physics 2015-08-11 Cristiano Roberto Fabri Granzotti , Alexandre Souto Martinez

A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…

Machine Learning · Computer Science 2022-10-18 Soumita Modak

Transportation distance information is a powerful resource, but location records are often censored due to privacy concerns or regulatory mandates. We outline methods to approximate, sample from, and compare distributions of distances…

Methodology · Statistics 2024-08-05 Lucas H. McCabe

Motivated by putting empirical work based on (synthetic) election data on a more solid mathematical basis, we analyze six distances among elections, including, e.g., the challenging-to-compute but very precise swap distance and the distance…

Computer Science and Game Theory · Computer Science 2022-05-03 Niclas Boehmer , Piotr Faliszewski , Rolf Niedermeier , Stanisław Szufa , Tomasz Wąs

Distance plays a fundamental role in measuring similarity between objects. Various visualization techniques and learning tasks in statistics and machine learning such as shape matching, classification, dimension reduction and clustering…

Machine Learning · Statistics 2025-04-23 Dianbin Bao , Kisung You , Lizhen Lin

Testing the independence between random vectors is a fundamental problem in statistics. Distance correlation, a recently popular dependence measure, is universally consistent for testing independence against all distributions with finite…

Methodology · Statistics 2024-08-22 Yuwei Ke , Hok Kan Ling , Yanglei Song

Finite mixture models that allow for a broad range of potentially non-elliptical cluster distributions is an emerging methodological field. Such methods allow for the shape of the clusters to match the natural heterogeneity of the data,…

Distance measures have been recognized as one of the fundamental building blocks in time-series analysis tasks, e.g., querying, indexing, classification, clustering, anomaly detection, and similarity search. The vast proliferation of…

Databases · Computer Science 2024-12-31 John Paparrizos , Haojun Li , Fan Yang , Kaize Wu , Jens E. d'Hondt , Odysseas Papapetrou

The notions of distance and similarity play a key role in many machine learning approaches, and artificial intelligence (AI) in general, since they can serve as an organizing principle by which individuals classify objects, form concepts…

Artificial Intelligence · Computer Science 2020-02-19 Santiago Ontañón

To quantify the fundamental evolution of time-varying networks, and detect abnormal behavior, one needs a notion of temporal difference that captures significant organizational changes between two successive instants. In this work, we…

Social and Information Networks · Computer Science 2017-08-17 Nathan D Monnig , Francois G Meyer

Clustering in high dimension spaces is a difficult task; the usual distance metrics may no longer be appropriate under the curse of dimensionality. Indeed, the choice of the metric is crucial, and it is highly dependent on the dataset…

Machine Learning · Computer Science 2023-02-14 Simo Alami. C , Rim Kaddah , Jesse Read

Graphs are used in almost every scientific discipline to express relations among a set of objects. Algorithms that compare graphs, and output a closeness score, or a correspondence among their nodes, are thus extremely important. Despite…

Discrete Mathematics · Computer Science 2020-11-17 Sam Safavi , José Bento

The need for appropriate ways to measure the distance or similarity between data is ubiquitous in machine learning, pattern recognition and data mining, but handcrafting such good metrics for specific problems is generally difficult. This…

Machine Learning · Computer Science 2019-01-25 Aurélien Bellet , Amaury Habrard , Marc Sebban

Given two distinct datasets, an important question is if they have arisen from the the same data generating function or alternatively how their data generating functions diverge from one another. In this paper, we introduce an approach for…

Machine Learning · Statistics 2019-09-17 Marco Henrique de Almeida Inácio , Rafael Izbicki , Bálint Gyires-Tóth

Acyclic digraphs arise in many natural and artificial processes. Among the broader set, dynamic citation networks represent a substantively important form of acyclic digraphs. For example, the study of such networks includes the spread of…

Physics and Society · Physics 2011-07-26 Michael J. Bommarito , Daniel Martin Katz , Jon Zelner , James H. Fowler

Clustering is an important part of many modern data analysis pipelines, including network analysis and data retrieval. There are many different clustering algorithms developed by various communities, and it is often not clear which…

Machine Learning · Computer Science 2019-10-04 Maria-Florina Balcan , Travis Dick , Manuel Lang

Data-dependent metrics are powerful tools for learning the underlying structure of high-dimensional data. This article develops and analyzes a data-dependent metric known as diffusion state distance (DSD), which compares points using a…

Machine Learning · Statistics 2020-03-10 Lenore Cowen , Kapil Devkota , Xiaozhe Hu , James M. Murphy , Kaiyi Wu