English
Related papers

Related papers: Surprises in approximating Levenshtein distances

200 papers

This paper hypothesizes that chunking plays important role in reducing dependency distance and dependency crossings. Computer simulations, when compared with natural languages,show that chunking reduces mean dependency distance (MDD) of a…

Computation and Language · Computer Science 2016-09-27 Qian Lu , Chunshan Xu , Haitao Liu

Modern neural sequence generation models are built to either generate tokens step-by-step from scratch or (iteratively) modify a sequence of tokens bounded by a fixed length. In this work, we develop Levenshtein Transformer, a new partially…

Computation and Language · Computer Science 2019-10-29 Jiatao Gu , Changhan Wang , Jake Zhao

Important data mining problems such as nearest-neighbor search and clustering admit theoretical guarantees when restricted to objects embedded in a metric space. Graphs are ubiquitous, and clustering and classification over graphs arise in…

Combinatorics · Mathematics 2018-01-16 Jose Bento , Stratis Ioannidis

To what extent can we distinguish one probability distribution from another? Are there quantitative measures of distinguishability? The goal of this tutorial is to approach such questions by introducing the notion of the "distance" between…

Data Analysis, Statistics and Probability · Physics 2015-06-23 Ariel Caticha

In this article, we generalize the Wasserstein distance to measures with different masses. We study the properties of such distance. In particular, we show that it metrizes weak convergence for tight sequences. We use this generalized…

Analysis of PDEs · Mathematics 2015-06-05 Benedetto Piccoli , Francesco Rossi

The primary research questions of this paper center on defining the amount of context that is necessary and/or appropriate when investigating the relationship between language model probabilities and cognitive phenomena. We investigate…

Computation and Language · Computer Science 2026-01-07 Cassandra L. Jacobs , Andrés Buxó-Lugo , Anna K. Taylor , Marie Leopold-Hooke

Generalized Wasserstein distances allow to quantitatively compare two continuous or atomic mass distributions with equal or different total mass. In this paper, we propose four numerical methods for the approximation of three different…

Analysis of PDEs · Mathematics 2025-10-21 Maya Briani , Emiliano Cristiani , Giovanni Franzina , Francesca L. Ignoto

Word embeddings are high dimensional vector representations of words that capture their semantic similarity in the vector space. There exist several algorithms for learning such embeddings both for a single language as well as for several…

Computation and Language · Computer Science 2019-11-12 Georgios Balikas , Ioannis Partalas

Measuring similarity is a basic task in information retrieval, and now often a building-block for more complex arguments about cultural change. But do measures of textual similarity and distance really correspond to evidence about cultural…

Computation and Language · Computer Science 2018-07-03 Ted Underwood

By a method inspired of the Stein's method, we derive an upper-bound of the Rubinstein distance between two absolutely continuous probability measures on configurations space. As an application, we show that the best way to approximate a…

Probability · Mathematics 2007-07-04 Laurent Decreusefond , Nicolas Savy

Statistical distances, divergences, and similar quantities have a large history and play a fundamental role in statistics, machine learning and associated scientific disciplines. However, within the statistical literature, this extensive…

Statistics Theory · Mathematics 2018-06-08 Marianthi Markatou , Yang Chen , Georgios Afendras , Bruce G. Lindsay

Similarity measures are used extensively in machine learning and data science algorithms. The newly proposed graph Relative Hausdorff (RH) distance is a lightweight yet nuanced similarity measure for quantifying the closeness of two graphs.…

Discrete Mathematics · Computer Science 2019-06-13 Sinan G. Aksoy , Kathleen E. Nowak , Emilie Purvine , Stephen J. Young

Approximate Bayesian Computation (ABC) is a popular method for approximate inference in generative models with intractable but easy-to-sample likelihood. It constructs an approximate posterior distribution by finding parameters for which…

Computation · Statistics 2020-03-09 Kimia Nadjahi , Valentin De Bortoli , Alain Durmus , Roland Badeau , Umut Şimşekli

Graphs are used in many disciplines to model the relationships that exist between objects in a complex discrete system. Researchers may wish to compare a network of interest to a "typical" graph from a family (or ensemble) of graphs which…

Combinatorics · Mathematics 2025-08-08 Catherine Greenhill

Stochastic approximation algorithm is a useful technique which has been exploited successfully in probability theory and statistics for a long time. The step sizes used in stochastic approximation are generally taken to be deterministic and…

Probability · Mathematics 2019-09-25 Ujan Gangopadhyay , Krishanu Maulik

There is evidence that the numbers in probabilistic inference don't really matter. This paper considers the idea that we can make a probabilistic model simpler by making fewer distinctions. Unfortunately, the level of a Bayesian network…

Artificial Intelligence · Computer Science 2013-02-01 David L. Poole

An approximation method is presented for probabilistic inference with continuous random variables. These problems can arise in many practical problems, in particular where there are "second order" probabilities. The approximation, based on…

Artificial Intelligence · Computer Science 2013-04-10 Ross D. Shachter

The Sliced-Wasserstein distance (SW) is being increasingly used in machine learning applications as an alternative to the Wasserstein distance and offers significant computational and statistical benefits. Since it is defined as an…

Machine Learning · Statistics 2022-01-05 Kimia Nadjahi , Alain Durmus , Pierre E. Jacob , Roland Badeau , Umut Şimşekli

The Fr\'echet distance is a popular distance measure for curves which naturally lends itself to fundamental computational tasks, such as clustering, nearest-neighbor searching, and spherical range searching in the corresponding metric…

Computational Geometry · Computer Science 2018-08-07 Anne Driemel , Amer Krivošija

Distance covariance is a popular measure of dependence between random variables. It has some robustness properties, but not all. We prove that the influence function of the usual distance covariance is bounded, but that its breakdown value…

Methodology · Statistics 2025-08-26 Sarah Leyder , Jakob Raymaekers , Peter J. Rousseeuw