English
Related papers

Related papers: Surprises in approximating Levenshtein distances

200 papers

Several important algorithms for machine learning and data analysis use pairwise distances as input. On Riemannian manifolds these distances may be prohibitively costly to compute, in particular for large datasets. To tackle this problem,…

Differential Geometry · Mathematics 2019-04-29 Philipp Harms , Elodie Maignant , Stefan Schlager

In this paper a notion of functional "distance" in the Mellin transform setting is introduced and a general representation formula is obtained for it. Also, a determination of the distance is given in terms of Lipschitz classes and…

Functional Analysis · Mathematics 2016-03-15 Carlo Bardaro , Paul L. Butzer , Ilaria Mantellini , Gerhard Schmeisser

This document reviews the definition of the kernel distance, providing a gentle introduction tailored to a reader with background in theoretical computer science, but limited exposure to technology more common to machine learning,…

Computational Geometry · Computer Science 2011-03-11 Jeff M. Phillips , Suresh Venkatasubramanian

In this paper we examine the usefulness of two classes of algorithms Distance Methods, Discrete Character Methods (Felsenstein and Felsenstein 2003) widely used in genetics, for predicting the family relationships among a set of related…

Computation and Language · Computer Science 2014-01-06 Taraka Rama , Sudheer Kolachina , Lakshmi Bai B

Given an arbitrary long but finite sequence of observations from a finite set, we construct a simple process that approximates the sequence, in the sense that with high probability the empirical frequency, as well as the empirical one-step…

Statistics Theory · Mathematics 2007-06-13 Dinah Rosenberg , Eilon Solan , Nicolas Vieille

This research project aimed to overcome the challenge of analysing human language relationships, facilitate the grouping of languages and formation of genealogical relationship between them by developing automated comparison techniques.…

Computation and Language · Computer Science 2020-02-03 Gabija Mikulyte , David Gilbert

Storing information in DNA molecules is of great interest because of its advantages in longevity, high storage density, and low maintenance cost. A key step in the DNA storage pipeline is to efficiently cluster the retrieved DNA sequences…

Machine Learning · Computer Science 2022-07-12 Alan J. X. Guo , Cong Liang , Qing-Hu Hou

The ability to mimic human notions of semantic distance has widespread applications. Some measures rely only on raw text (distributional measures) and some rely on knowledge sources such as WordNet. Although extensive studies have been…

Computation and Language · Computer Science 2012-03-09 Saif M. Mohammad , Graeme Hirst

Wasserstein distances provide a powerful framework for comparing data distributions. They can be used to analyze processes over time or to detect inhomogeneities within data. However, simply calculating the Wasserstein distance or analyzing…

Machine Learning · Computer Science 2026-03-03 Philip Naumann , Jacob Kauffmann , Grégoire Montavon

A new numerical characterization of symbolic sequences is proposed. The partition of sequence based on Ke and Tong algorithm is a starting point. Algorithm decomposes original sequence into set of distinct subsequences - a patterns. The set…

Quantitative Methods · Quantitative Biology 2011-09-08 B. Kozarzewski

We define a (pseudo-)distance between graphs based on the spectrum of the normalized Laplacian, which is easy to compute or to estimate numerically. It can therefore serve as a rough classification of large empirical graphs into families…

Spectral Theory · Mathematics 2019-04-03 Jiao Gu , Jürgen Jost , Shiping Liu , Peter F. Stadler

Statistical models often include thousands of parameters. However, large models decrease the investigator's ability to interpret and communicate the estimated parameters. Reducing the dimensionality of the parameter space in the estimation…

Methodology · Statistics 2022-05-16 Eric Dunipace , Lorenzo Trippa

The nested distance builds on the Wasserstein distance to quantify the difference of stochastic processes, including also the information modelled by filtrations. The Sinkhorn divergence is a relaxation of the Wasserstein distance, which…

Optimization and Control · Mathematics 2021-02-11 Alois Pichler , Michael Weinhardt

Chamfer distances play an important role in the theory of distance transforms. Though the determination of the exact Euclidean distance transform is also a well investigated area, the classical chamfering method based upon "small"…

Information Theory · Computer Science 2012-01-05 Andras Hajdu , Lajos Hajdu , Robert Tijdeman

We develop a projected Wasserstein distance for the two-sample test, a fundamental problem in statistics and machine learning: given two sets of samples, to determine whether they are from the same distribution. In particular, we aim to…

Machine Learning · Statistics 2024-04-01 Jie Wang , Rui Gao , Yao Xie

Many functions have been recently defined to assess the similarity among networks as tools for quantitative comparison. They stem from very different frameworks - and they are tuned for dealing with different situations. Here we show an…

Molecular Networks · Quantitative Biology 2012-08-21 Giuseppe Jurman , Roberto Visintainer , Cesare Furlanello

This paper proposes a general framework for matching similar subsequences in both time series and string databases. The matching results are pairs of query subsequences and database subsequences. The framework finds all possible pairs of…

Databases · Computer Science 2012-08-02 Haohan Zhu , George Kollios , Vassilis Athitsos

The goal of this thesis is to study the use of the Kantorovich-Rubinstein distance as to build a descriptor of sample complexity in classification problems. The idea is to use the fact that the Kantorovich-Rubinstein distance is a metric in…

Probability · Mathematics 2023-09-19 Gaël Giordano

Languages evolve over time in a process in which reproduction, mutation and extinction are all possible, similar to what happens to living organisms. Using this similarity it is possible, in principle, to build family trees which show the…

Computation and Language · Computer Science 2012-07-03 Maurizio Serva

Recent works on word representations mostly rely on predictive models. Distributed word representations (aka word embeddings) are trained to optimally predict the contexts in which the corresponding words tend to appear. Such models have…

Computation and Language · Computer Science 2015-04-10 Rémi Lebret , Ronan Collobert