基于Burrows Wheeler变换的云端序列索引大数据方法
分布式、并行与集群计算
2020-07-21 v1 人工智能
数据结构与算法
摘要
在精准医学背景下,序列数据索引十分重要,因为每天必须收集和分析大量“组学”数据,以便对患者进行分类并确定最有效的疗法。本文提出一种基于大数据技术(即Apache Spark和Hadoop)计算Burrows Wheeler变换的算法。我们的方法是首个将索引计算而非仅输入数据集进行分布的方法,从而能够充分利用可用的云资源。
引用
@article{arxiv.2007.10095,
title = {A Big Data Approach for Sequences Indexing on the Cloud via Burrows Wheeler Transform},
author = {Mario Randazzo and Simona E. Rombo},
journal= {arXiv preprint arXiv:2007.10095},
year = {2020}
}
备注
Accepted at HELPLINE@ECAI2020