English

Spaced seeds improve k-mer-based metagenomic classification

Genomics 2016-03-17 v3 Computational Engineering, Finance, and Science Machine Learning

Abstract

Metagenomics is a powerful approach to study genetic content of environmental samples that has been strongly promoted by NGS technologies. To cope with massive data involved in modern metagenomic projects, recent tools [4, 39] rely on the analysis of k-mers shared between the read to be classified and sampled reference genomes. Within this general framework, we show in this work that spaced seeds provide a significant improvement of classification accuracy as opposed to traditional contiguous k-mers. We support this thesis through a series a different computational experiments, including simulations of large-scale metagenomic projects. Scripts and programs used in this study, as well as supplementary material, are available from http://github.com/gregorykucherov/spaced-seeds-for-metagenomics.

Keywords

Cite

@article{arxiv.1502.06256,
  title  = {Spaced seeds improve k-mer-based metagenomic classification},
  author = {Karel Brinda and Maciej Sykulski and Gregory Kucherov},
  journal= {arXiv preprint arXiv:1502.06256},
  year   = {2016}
}

Comments

23 pages

R2 v1 2026-06-22T08:34:58.276Z