English

Scalable Probabilistic Entity-Topic Modeling

Machine Learning 2013-09-03 v1 Information Retrieval Machine Learning

Abstract

We present an LDA approach to entity disambiguation. Each topic is associated with a Wikipedia article and topics generate either content words or entity mentions. Training such models is challenging because of the topic and vocabulary size, both in the millions. We tackle these problems using a novel distributed inference and representation framework based on a parallel Gibbs sampler guided by the Wikipedia link graph, and pipelines of MapReduce allowing fast and memory-frugal processing of large datasets. We report state-of-the-art performance on a public dataset.

Keywords

Cite

@article{arxiv.1309.0337,
  title  = {Scalable Probabilistic Entity-Topic Modeling},
  author = {Neil Houlsby and Massimiliano Ciaramita},
  journal= {arXiv preprint arXiv:1309.0337},
  year   = {2013}
}
R2 v1 2026-06-22T01:18:56.235Z