两个随机序列间的近似词匹配
概率论
2009-09-29 v1
摘要
给定有限字母表 上的两个序列, 统计量是这两个序列之间 字母词匹配的数量。该统计量用于生物信息学中的表达序列标签数据库搜索。在此,我们在 DNA 序列的背景下研究 统计量的推广,假设文本具有链对称伯努利分布。对于 ,我们考察最多允许 个错配的 字母词匹配计数。针对该统计量,我们计算了其期望值,给出了方差的上下界,并证明了其分布是渐近正态的。
引用
@article{arxiv.0801.3145,
title = {Approximate word matches between two random sequences},
author = {Conrad J. Burden and Miriam R. Kantorovitz and Susan R. Wilson},
journal= {arXiv preprint arXiv:0801.3145},
year = {2009}
}
备注
Published in at http://dx.doi.org/10.1214/07-AAP452 the Annals of Applied Probability (http://www.imstat.org/aap/) by the Institute of Mathematical Statistics (http://www.imstat.org)