English

WordRank: Learning Word Embeddings via Robust Ranking

Computation and Language 2016-09-29 v4 Machine Learning Machine Learning

Abstract

Embedding words in a vector space has gained a lot of attention in recent years. While state-of-the-art methods provide efficient computation of word similarities via a low-dimensional matrix embedding, their motivation is often left unclear. In this paper, we argue that word embedding can be naturally viewed as a ranking problem due to the ranking nature of the evaluation metrics. Then, based on this insight, we propose a novel framework WordRank that efficiently estimates word representations via robust ranking, in which the attention mechanism and robustness to noise are readily achieved via the DCG-like ranking losses. The performance of WordRank is measured in word similarity and word analogy benchmarks, and the results are compared to the state-of-the-art word embedding techniques. Our algorithm is very competitive to the state-of-the- arts on large corpora, while outperforms them by a significant margin when the training set is limited (i.e., sparse and noisy). With 17 million tokens, WordRank performs almost as well as existing methods using 7.2 billion tokens on a popular word similarity benchmark. Our multi-node distributed implementation of WordRank is publicly available for general usage.

Keywords

Cite

@article{arxiv.1506.02761,
  title  = {WordRank: Learning Word Embeddings via Robust Ranking},
  author = {Shihao Ji and Hyokun Yun and Pinar Yanardag and Shin Matsushima and S. V. N. Vishwanathan},
  journal= {arXiv preprint arXiv:1506.02761},
  year   = {2016}
}

Comments

Conference on Empirical Methods in Natural Language Processing (EMNLP), November 1-5, 2016, Austin, Texas, USA

R2 v1 2026-06-22T09:49:49.358Z