Surprises in approximating Levenshtein distances
Quantitative Methods
2007-05-23 v2
Abstract
The Levenshtein distance is an important tool for the comparison of symbolic sequences, with many appearances in genome research, linguistics and other areas. For efficient applications, an approximation by a distance of smaller computational complexity is highly desirable. However, our comparison of the Levenshtein with a generic dictionary-based distance indicates their statistical independence. This suggests that a simplification along this line might not be possible without restricting the class of sequences. Several other probabilistic properties are briefly discussed, emphasizing various questions that deserve further investigation.
Cite
@article{arxiv.q-bio/0601006,
title = {Surprises in approximating Levenshtein distances},
author = {Michael Baake and Uwe Grimm and Robert Giegerich},
journal= {arXiv preprint arXiv:q-bio/0601006},
year = {2007}
}
Comments
7 pages, 4 figures; revised version