English

NASCUP: Nucleic Acid Sequence Classification by Universal Probability

Genomics 2018-11-30 v2 Information Theory math.IT

Abstract

Motivated by the need for fast and accurate classification of unlabeled nucleotide sequences on a large scale, we developed NASCUP, a new classification method that captures statistical structures of nucleotide sequences by compact context-tree models and universal probability from information theory. NASCUP achieved BLAST-like classification accuracy consistently for several large-scale databases in orders-of-magnitude reduced runtime, and was applied to other bioinformatics tasks such as outlier detection and synthetic sequence generation.

Keywords

Cite

@article{arxiv.1511.04944,
  title  = {NASCUP: Nucleic Acid Sequence Classification by Universal Probability},
  author = {Sunyoung Kwon and Gyuwan Kim and Byunghan Lee and Jongsik Chun and Sungroh Yoon and Young-Han Kim},
  journal= {arXiv preprint arXiv:1511.04944},
  year   = {2018}
}
R2 v1 2026-06-22T11:46:12.815Z