中文

两个随机序列间的近似词匹配

概率论 2009-09-29 v1

摘要

给定有限字母表 L\mathcal{L} 上的两个序列,D2D_2 统计量是这两个序列之间 mm 字母词匹配的数量。该统计量用于生物信息学中的表达序列标签数据库搜索。在此,我们在 DNA 序列的背景下研究 D2D_2 统计量的推广,假设文本具有链对称伯努利分布。对于 k<mk<m,我们考察最多允许 kk 个错配的 mm 字母词匹配计数。针对该统计量,我们计算了其期望值,给出了方差的上下界,并证明了其分布是渐近正态的。

关键词

引用

@article{arxiv.0801.3145,
  title  = {Approximate word matches between two random sequences},
  author = {Conrad J. Burden and Miriam R. Kantorovitz and Susan R. Wilson},
  journal= {arXiv preprint arXiv:0801.3145},
  year   = {2009}
}

备注

Published in at http://dx.doi.org/10.1214/07-AAP452 the Annals of Applied Probability (http://www.imstat.org/aap/) by the Institute of Mathematical Statistics (http://www.imstat.org)