中文

连接的读取策略:渐近分析

概率论 2007-05-23 v1 数据库

摘要

假设从分布 R\mathbf {R} 中抽取 mnm_n 个观测,从分布 S\mathbf {S} 中抽取 nmnn-m_n 个观测。对每一对来自 R\mathbf {R}xx 和来自 S\mathbf {S}yy,赋予一个非负得分 ϕ(x,y)\phi(x,y)。最优读取策略是产生序列 mnm_n 从而一致地关于 nn 最大化 E(M(n))\mathbb{E}(M(n))(即 (nmn)mn(n-m_n)m_n 个观测得分的期望和)的策略。交替策略(在两个来源间切换)是最优非自适应策略。相反,贪婪策略(选择其来源以最大化下一步的期望增益)被证明是最优策略。对 R\mathbf {R}S\mathbf {S} 分布为离散且 ϕ(x,y)=1\phi(x,y)=100(依 x=yx=y 与否,即观测是否匹配)的情形给出渐近性。具体地,证明了一个不变性结果,保证对包括交替与贪婪在内的广泛策略类,变量 M(n) 服从相同的 CLT 与 LIL。对交替与贪婪两者的序列 E(M(n))\mathbb{E}(M(n)) 和 M(n) 样本路径的更精细分析,揭示了后者策略在渐近意义上优于前者的微弱程度,以及两者等价性的一种意义和前者的鲁棒性。

关键词

引用

@article{arxiv.math/0703019,
  title  = {Reading policies for joins: An asymptotic analysis},
  author = {Ralph P. Russo and Nariankadu D. Shyamalkumar},
  journal= {arXiv preprint arXiv:math/0703019},
  year   = {2007}
}

备注

Published at http://dx.doi.org/10.1214/105051606000000646 in the Annals of Applied Probability (http://www.imstat.org/aap/) by the Institute of Mathematical Statistics (http://www.imstat.org)