松弛可辨识条件下潜话题模型的收敛速率
机器学习
2019-01-21 v2 机器学习
摘要
本文研究了潜狄利克雷分配(Blei 等,2003)话题模型的频率派收敛速率。我们证明了极大似然估计量在 Wasserstein 距离度量下以 的速率收敛到有限多个等价参数之一,且无需假设底层话题的可分离性或非退化性,也无需假设每篇文档存在多于三个词,从而从信息论角度推广了 Anandkumar 等(2012, 2014)的先前工作。我们还证明了 的收敛速率在最坏情况下是最优的。
引用
@article{arxiv.1710.11070,
title = {Convergence Rates of Latent Topic Models Under Relaxed Identifiability Conditions},
author = {Yining Wang},
journal= {arXiv preprint arXiv:1710.11070},
year = {2019}
}
备注
26 pages, 1 table. Added significantly more expositions, and a numerical procedure to check the order of degeneracy. Proofs slightly altered with explicit constants given at various places