English

Multi-SpaM: a Maximum-Likelihood approach to Phylogeny reconstruction based on Multiple Spaced-Word Matches

Populations and Evolution 2018-05-01 v2

Abstract

Motivation: Word-based or `alignment-free' methods for phylogeny reconstruction are much faster than traditional approaches, but they are generally less accurate. Most of these methods calculate pairwise distances for a set of input sequences, for example from word frequencies, from so-called spaced-word matches or from the average length of common substrings. Results: In this paper, we propose the first word-based approach to tree reconstruction that is based on multiple sequence comparison and Maximum Likelihood. Our algorithm first samples small, gap-free alignments involving four taxa each. For each of these alignments, it then calculates a quartet tree and, finally, the program Quartet MaxCut is used to infer a super tree topology for the full set of input taxa from the calculated quartet trees. Experimental results show that trees calculated with our approach are of high quality. Availability: The source code of the program is available at https://github.com/tdencker/multi-SpaM Contact: thomas.dencker@stud.uni-goettingen.de

Keywords

Cite

@article{arxiv.1803.09222,
  title  = {Multi-SpaM: a Maximum-Likelihood approach to Phylogeny reconstruction based on Multiple Spaced-Word Matches},
  author = {Thomas Dencker and Chris-Andre Leimeister and Michael Gerth and Christoph Bleidorn and Sagi Snir and Burkhard Morgenstern},
  journal= {arXiv preprint arXiv:1803.09222},
  year   = {2018}
}
R2 v1 2026-06-23T01:04:12.602Z