基于窗口采样与负采样的矩阵分解以改进词表示
计算与语言
2016-06-08 v2
摘要
本文提出 LexVec,一种生成分布式词表示的新方法。该方法通过随机梯度下降对正点互信息矩阵进行低秩加权分解,采用一种加权方案,对频繁共现上的误差施加更重的惩罚,同时仍考虑负共现。在词语相似度和类比任务上的评估表明,LexVec 在这些任务中的许多项上匹敌并常常超越当前最优方法。
引用
@article{arxiv.1606.00819,
title = {Matrix Factorization using Window Sampling and Negative Sampling for Improved Word Representations},
author = {Alexandre Salle and Marco Idiart and Aline Villavicencio},
journal= {arXiv preprint arXiv:1606.00819},
year = {2016}
}
备注
Converted paper size from A4 to US Letter to avoid margin issues on arXiv